New York City · Senior Software Engineer
Backend systems
for the AI era.
I'm Manan Chawda. I build production infrastructure where scale, reliability, and AI meet: high-throughput backend platforms, cloud data pipelines, retrieval systems, and the observability that keeps them honest.
The throughline in my work is building systems that behave under pressure: payment streams, advertising data platforms, cloud migrations, and LLM workflows that need evaluation instead of vibes.
AdTech
Experian
FinTech
Barclays
Backend · Cloud · Data pipelines · RAG · Agents · Observability · CI/CD
6
years building production systems
100TB+
cloud data estate migrated
300M+
consumer records served
10M+
daily payment transactions
<2%
RAG hallucination rate
What I should be known for
Experienced enough to ship the system. Curious enough to rebuild the hard parts.
Your brand should not be "AI engineer" in the generic sense. The stronger story is rarer: a backend and cloud engineer who has already operated at serious scale, now applying that discipline to RAG, agents, evaluation, and AI reliability.
Backend platforms
Event-driven services, Kafka systems, CQRS read paths, gRPC APIs, and migration work where correctness matters.
Cloud & data infrastructure
AWS, EMR, Spark, Airflow, Terraform, Kubernetes, and observability for systems operating at 100TB+ and 300M+ record scale.
AI infrastructure
Production RAG, LangGraph agents, hybrid retrieval, LLM-as-judge evaluation, regression testing, and cost-aware serving.
Proof Built under real load
full experience →Experian
Modernized audience data and AI operations
Owned ConsumerView pipeline work across 300M+ records and 100TB+ on AWS, rebuilt batch workloads on Spark/EMR, and shipped a multi-agent RAG assistant that turned operational investigations from 30 minutes into sub-5-second answers.
99.5% DAG SLA
Barclays
Rebuilt payments and rewards systems for real-time reliability
Moved rewards from nightly batch to Kafka streaming with exactly-once guarantees, built event-sourced spend trackers, and led Kubernetes migrations for card platforms serving millions of users.
10M+ txns/day
Open-source lab
Turns architecture curiosity into working systems
Builds practical references for vector ingestion, RAG regression testing, distributed queues, Raft-style coordination, and failure-mode testing instead of stopping at diagrams.
500 docs/sec
Portfolio Selected systems
all projects →Real-Time Vector Ingestion Engine
Streaming RAG ingestion pipeline with Kafka backpressure, batched embeddings, provider fallback, idempotent pgvector upserts, DLQ retries, and p95 freshness under 10 seconds.
LLM Evaluation Harness for RAG Regression Testing
Open-source CI framework for RAG regression with golden sets, LLM-as-judge scoring, GitHub Actions merge gates, Prometheus drift tracking, and sub-5-second p95 eval runs.
Multi-Agent RAG Orchestration
A 4-agent retrieval-augmented system on LangGraph with RAGAS evaluation. Faithfulness 73% → 91% over a single-chain baseline.
Writing How I think
all posts →- 01
Run the RAG eval before you redesign the architecture
A production RAG lesson: evaluation should tell you whether to add agents, not become the thing you bolt on after the graph already exists.
- 02
Real-time vector ingestion is a data platform problem
Fresh retrieval depends on unglamorous infrastructure: Kafka backpressure, idempotent upserts, embedding retries, DLQs, and freshness metrics.
- 03
What batch-to-streaming migrations teach you about reliability
Lessons from moving rewards and data workloads from scheduled batch jobs toward streaming systems with clearer ownership, latency, and failure boundaries.
- 04
Multi-Agent RAG with LangGraph: when 4 agents beat 1
Building a production retrieval-augmented system with router, retriever, verifier, and synthesizer agents — and what changed when we added RAGAS evaluation.
Next role
I am looking for teams building serious backend, cloud, data, or AI infrastructure.
The work I want is the kind where architecture, ownership, and production judgment all matter.