Samir Sawarkar
Engineering reliability. Building deterministic safety gates, evaluation harnesses, and trace/replay infrastructure for AI systems touching systems of record.
In the past I’ve built manufacturing BOM validation pipelines (EBOM → MBOM), financial ERP invoice verification gates, and operated platforms for 50,000+ users.
Building
- FAULTLINE
Reliability workbench for document-grounded AI.
- SafeGate (EBOM) ↗
Deterministic validation pipeline with zero silent corruption.
- Reliability Lab ✦
Gated harness, fault taxonomy, and evaluation metrics.
Projects
- FAULTLINE ↗
Open source failure injection and replay harness.
- Invoice → ERP Gate
Deterministic safety gate for financial ledger writes.
- Research Paper ↗
Published study on ERP manufacturing validation.
- All projects
Archive of systems, tools, and production gates.
Writing
- GPT-6 Astra Scored 100%. Another Test Gave It 39%. Which Number Should You Trust?
AI scores keep climbing. Here is why the number on the slide and the number in your business are never the same one — and what that gap costs.
- Why AI Gets Worse the Longer It Works
AI agents do not get tired. They get convinced—and that quiet failure mode is becoming the reliability tax of long-running AI systems.
- The Last 3% Is Invisible Until You Chain It
Why you can't tell two models apart in chat — and why that gap decides whether your agent finishes its task or dies quietly at step fourteen.
- All writing
Infrequent thoughts on reliability, gates, and code.
Now
Based in Pune, India. Deep in deterministic safety contracts for multi-agent LLM systems, fault injection benchmarks, and statistical evaluation boundaries. Mindful that everything around me is someone’s life work.
All I want to do is build reliable software. Deterministic verification, trace reconstruction, evaluation harnesses, and systems that fail loudly before production.
Connect
Reach me on X or by email. Code on GitHub, research on ResearchGate.