Software Development Engineer I – Data Engineering
- Built real-time event collection and data ingestion components on a highly scalable pipeline: PostgreSQL change data captured via Debezium CDC into Apache Kafka, processed with Apache Spark (Structured Streaming) and Apache Flink across batch and streaming workloads.
- Developed data pipelines supporting AI model training and autonomous AI agents — feature engineering, embedding generation, and vector retrieval (Milvus) — and operated MCP servers exposing governed datasets and tools for LLM access.
- Implemented ETL/ELT pipelines orchestrated with Apache Airflow over Delta Lake and Databricks, writing clean, well-tested Python and SQL under code review, Git, and CI/CD standards.
- Contributed to data governance, observability, and quality; identified and implemented performance optimizations for ingestion, storage, and processing components in production.
