AboutI design, build, and operate the full ML lifecycle — from data engineering through production deployment, monitoring, and cost optimization.
My work centers on intelligent systems that hold up in production: multi-agent workflows, retrieval-augmented generation, and distributed inference that stay low-latency and reliable under enterprise load. I care as much about the observability dashboard and the GPU bill as I do about model accuracy.
GenAILLMs, RAG, multi-agent and agentic systems, fine-tuning, guardrails
ServingvLLM, Ray Serve, model routing, dynamic batching, GPU optimization
MLOpsMLflow 3.0, CI/CD, evaluation and monitoring pipelines
PlatformKubernetes, AWS, FastAPI microservices, event-driven orchestration