Skip to content

I spent nine years making systems not fall over. Now I do it for agents.

I am a technical lead based in Pune, India, and my background is deliberately unfashionable for someone working in AI: distributed systems, data pipelines, and production support. I learned what reliability means by being the person who got paged.

The largest system I have architected processes 520 million telecom parameters every fifteen minutes on Golang, Kafka, Kubernetes and PostgreSQL. Not because throughput is impressive on its own, but because a fixed cycle time teaches you something a demo never will: a system that accepts more work than it can finish does not slow down, it collapses. That lesson turns out to be the single most useful thing I know about AI agents.

Since 2022 I have been building AI products — conversational analytics over sensitive multi-tenant data, multi-agent systems on CrewAI and Model Context Protocol, text-to-SQL engines, retrieval pipelines, an agentic voice assistant, autonomous lead-generation agents. Some of it worked immediately. More of it worked after I stopped trusting the model with things a model should not be trusted with.

What I noticed across all of it: the teams struggling with agents were rarely struggling with AI. They were missing validation at boundaries, bounded resources, traceable failures, and any way to tell whether yesterday's change made things better. That is ordinary engineering discipline, and it iswhat I now do for a living — infour domains I have actually shipped in.

On working independently: this is a deliberate choice, not a gap between jobs. Reliability work is the part of AI engineering that consistently gets deprioritised inside a company, because features have an owner arguing for them and reliability does not. Working from outside means I get handed the problem directly, with a scope and a deadline, and I am not competing with a roadmap for attention. I take a small number of engagements at a time for the same reason.

I also work at the business end of things, which is unusual for an engineer and occasionally useful: I have led an initiative analysing customer data to find expansion opportunities, built AI systems aimed directly at pipeline and retention, and presented enterprise positioning strategy to a CEO. I mention it because it means I will ask what a reliability problem costs you, not just how to fix it.

Durgesh Rathod, AI reliability engineer based in Pune, India
Available for new work

Based in Pune, India (+05:30). Local time .

Your time
Working-hours overlap

I hold afternoons and evenings IST open for calls, which covers European mornings and US mornings on the earlier side. Async-first either way.

CV

Full history, projects and stack. Updated July 2026.

Download PDF

169 KB · PDF

Credentials

  • Red Hat Certified Engineer
  • Generative AI with Large Language Models
  • BE Computer Engineering, 2017

Published packages

The specifics

520M
params / 15 min

200M configuration + 320M performance parameters per 15-minute ingestion cycle, on Golang, Kafka, Kubernetes and PostgreSQL.

2,000
concurrent users

A text-to-SQL agent serving 2,000 concurrent users with strict per-tenant isolation on HR data.

30s
from 2–3 days

Conversational analytics replaced a manual reporting cycle that previously took two to three days.

9+
years shipping

Nine years across telecom, HR, recruitment and workflow automation — architecture through production support.

What I'm doing now

Updated

  • 1Taking on production readiness audits, and have room for one more retainer.
  • 2Published a failure taxonomy of nine named ways production agents break, with causes and fixes for each.
  • 3Writing about the cost factors most LLM calculators leave out — starting with tool definitions billed on every call.
  • 4Building a public benchmark comparing agent frameworks on identical tasks: tokens, cost, p95 latency and failure modes.

How I work

Worth reading before you hire anyone remote. These are commitments, not preferences.

Async-first, written down

You get a written findings document and a recorded walkthrough rather than a status meeting. Given the timezone gap that is not a compromise — it produces a better artefact, because writing forces the thinking to be finished.

Bad news arrives early

If an estimate is slipping or an approach is not working, you hear it from me while there is still time to change course. I would rather deliver an uncomfortable update in week one than a surprise in week three.

I will tell you not to hire me

If your gaps are ones your team can close, or the problem is not the one you think it is, I will say so. That has cost me engagements and earned me referrals, and I am not planning to stop.

You keep everything

Eval suites, traces, dashboards, documents. No dependency on me continuing, and nothing built in a way only I can maintain.

Don't take my word for it

Checking me out?

Nine years of history, recommendations and endorsements are on LinkedIn, and the code is on GitHub. For client references, ask me directly — I will put you in touch with people I have delivered for rather than sending you a curated quote.

Where I have worked

  1. Mar 2022 – present

    Technical Lead

    SourceFuse Technologies

    Architecture and delivery across AI products and large-scale data platforms. Built AI applications on OpenAI, Claude, AWS Bedrock, LangChain, CrewAI and Model Context Protocol, and architected a configuration-driven telecom platform processing 520M parameters per 15-minute cycle. Led design reviews, mentoring, release management and performance work across several teams.

  2. Feb 2021 – Feb 2022

    Senior Software Engineer

    SourceFuse Technologies

    Cloud-native backend services on AWS — microservices in Node.js and Python on ECS Fargate, CI/CD with CodePipeline, technical designs and mentoring.

  3. Dec 2018 – Jan 2021

    Software Engineer

    Flairlabs

    Backend services and enterprise applications on Node.js, AWS and SQL Server. REST APIs, database design, and production support across the full lifecycle.

  4. Jun 2017 – Dec 2018

    Application Engineer

    Merce Technologies

    Backend features and microservices, heavy SQL work, and first exposure to running things in production — which is where most of what I know actually came from.

Other things I have built

Not written up as case studies, but real and shipped. Happy to talk about any of them.

AI-powered recruitment platform

Candidate and job matching with NLP and vector embeddings, Elasticsearch search, resume parsing and ATS integrations, across employer and candidate portals.

PythonNode.jsElasticsearchBedrock

What makes this domain hard

Digital contract lifecycle management

Enterprise CLM covering authoring, review, approvals, amendments and clause libraries.

AngularNode.jsPostgreSQLCKEditor

What makes this domain hard

AI meeting assistant

Real-time Zoom audio transcription with Whisper, then summarisation and action-item extraction.

WhisperBedrockFastAPIDocker

Multi-agent long-form writing system

Specialised agents collaborating on story planning, character consistency, drafting and editorial review across a book-length manuscript. A serious exercise in long-horizon context management.

CrewAIMCPPython

Lost child detection (computer vision)

Facial recognition over CCTV footage, using augmentation to improve recognition on poor-quality frames.

PyTorchDlibPython

AI-driven GTM systems

Autonomous lead generation, WhatsApp retention workflows, and a HubSpot chatbot for qualification — AI applied to revenue rather than product features.

CrewAIHubSpotWhatsApp API

Want to talk?

About an agent that is misbehaving, a role you think fits, or a system you would like a second opinion on. I answer specific questions about specific systems for free.

Hiring rather than contracting?

I will listen to the right role. Send the company, the role and the comp range and I will be straight with you either way.

Download my CV