Open to AI/ML engineering roles

Arshitha
Ippagunta

AI/ML Engineer building production-scale Generative & Agentic AI.

Five years shipping LLM, RAG, and MLOps systems that stay fast, grounded, and cost-efficient at enterprise scale — from data pipeline to GPU inference to live monitoring.

5+
Years in production ML
8K+
Enterprise customers served
31%
Inference cost
29%
Response accuracy
About

I design, build, and operate the full ML lifecycle — from data engineering through production deployment, monitoring, and cost optimization.

My work centers on intelligent systems that hold up in production: multi-agent workflows, retrieval-augmented generation, and distributed inference that stay low-latency and reliable under enterprise load. I care as much about the observability dashboard and the GPU bill as I do about model accuracy.

GenAILLMs, RAG, multi-agent and agentic systems, fine-tuning, guardrails
ServingvLLM, Ray Serve, model routing, dynamic batching, GPU optimization
MLOpsMLflow 3.0, CI/CD, evaluation and monitoring pipelines
PlatformKubernetes, AWS, FastAPI microservices, event-driven orchestration
Experience

Where the systems ran

Two roles, one throughline: taking ML from notebook to a governed, monitored, cost-aware production service.

AI/ML Engineer

DatabricksSan Francisco, CA
Jun 2025 — Present
29% accuracy31% inference cost8,000+ customers
  • Architected scalable multi-agent AI workflows with Python, PyTorch, MLflow 3.0, and Ray, powering reliable enterprise task automation across production AI applications for 8,000+ customers.
  • Lifted response accuracy 29% via RAG optimization — hybrid retrieval, cross-encoder reranking, prompt engineering, and continuous evaluation on Vector Search and MLflow.
  • Cut inference cost 31% by tuning enterprise GPU serving with vLLM, Ray Serve, intelligent model routing, dynamic batching, and Kubernetes while maintaining low-latency performance.
  • Engineered LLM serving infrastructure and evaluation pipelines tracking response quality, hallucination detection, and safety guardrails.
  • Built cloud-native inference backends with FastAPI, Docker, Kubernetes, Helm, AWS EKS, and S3, plus event-driven orchestration with Ray, Redis, and PostgreSQL.
  • Operationalized end-to-end MLOps with MLflow 3.0, GitHub Actions, and Kubernetes, governed by Unity Catalog, Delta Lake, and Feature Store.

Machine Learning Engineer

AccentureIndia
Aug 2020 — Jul 2024
34% throughput40% deploy time12% accuracy
  • Engineered Spark-based feature engineering and model-training pipelines using PySpark, Spark ML, MLflow, and Delta Lake for consistent workflows across enterprise datasets.
  • Raised batch inference throughput 34% through optimized feature engineering, model caching, and Spark execution tuning.
  • Reduced end-to-end deployment time 40% with MLflow pipelines, Docker packaging, Kubernetes, Helm, and GitHub Actions CI/CD.
  • Improved production prediction accuracy 12% through feature engineering, automated hyperparameter optimization, and continuous evaluation.
  • Designed cloud-native ML infrastructure on AWS and shipped real-time and batch inference services using FastAPI and gRPC.
  • Established observability with Prometheus, Grafana, Fluent Bit, and MLflow metrics for reliable production operations.
Selected work

Two systems, measured

Retrieval and document intelligence platforms built to be grounded, fast, and evaluable.

Enterprise RAG Knowledge Assistant

A citation-backed question-answering platform over 50K+ enterprise documents — automated ingestion, embeddings, vector indexing, and evaluation around a tuned LLM inference stack.

35%
Retrieval precision
42%
Hallucinations
40%
Throughput
50K+
Documents
PythonFastAPILangChainLlamaIndexFAISSLlama 3MLflowKubernetes

AI Document Intelligence Platform

OCR-driven extraction, classification, and Q&A across 100K+ documents at 96% extraction accuracy — scalable FastAPI microservices with Redis caching and hybrid retrieval.

96%
Extraction accuracy
65%
Review time
38%
Search relevance
45%
Latency
PythonFastAPIHaystackLangChainFAISSRedisDockerAWS
Stack

Tools in rotation

Generative AI & LLMs

LLMsRAGAgentic AIMulti-Agent SystemsPrompt EngineeringFine-TuningEmbeddingsHybrid RetrievalCross-Encoder RerankingVector SearchGuardrailsModel Routing

ML & Serving

PyTorchScikit-learnSpark MLvLLMRayRay ServeMLflow 3.0Model RegistryGPU OptimizationHyperparameter TuningDeep Learning

Data & Backend

PythonSQLPySparkApache SparkAirflowKafkaDelta LakeDatabricksUnity CatalogFastAPIgRPCRedisPostgreSQL

Cloud, DevOps & Observability

AWSEKSEC2S3DockerKubernetesHelmTerraformGitHub ActionsCI/CDPrometheusGrafana
Certifications

Verified

GenAIDatabricks Certified Generative AI Engineer — Associate
MLAWS Certified Machine Learning Engineer — Associate
SAAWS Certified Solutions Architect — Associate
CKADCertified Kubernetes Application Developer
Education

Foundations

M.S. Computer Science & Engineering
University at Buffalo
B.Tech Computer Science & Engineering
Sree Vidyanikethan Engineering College
Contact

Let's build something
that ships and scales.

Open to AI/ML engineering roles and collaborations. The fastest way to reach me is email or LinkedIn.