Agentic AI & LLMs
Autonomous agents, multi-agent orchestration, RAG, and fine-tuning.
Samir Sengupta
Open to work

I frame the problem, build the model, and ship the system.
Multi-agent systems, RAG, and LLM pipelines, engineered onto Kubernetes with the data and evaluation infrastructure to keep them honest.
stack
intermission · the input device
Every pipeline on this page started at a keyboard. This one works: scroll lifts the keys off the plate, and typing plays it, sound on.

Most engineers pick a lane - research, modeling, or infrastructure. I work the whole line. Across 3+ years I’ve framed problems as a data scientist, built the models, and engineered them into reliable software: Kafka pipelines and FastAPI services on Kubernetes at sub-100ms latency, and predictive models that cut churn and lifted conversion for the teams that shipped them.
I’m the author of a research preprint on efficient long-context LLMs, spoke at KCD New York and the Apache Beam Summit in 2026, and hold an MS in Data Science (GPA 3.9) plus industry certifications from AWS, Google, NVIDIA, and IBM.
Autonomous agents, multi-agent orchestration, RAG, and fine-tuning.
Forecasting, anomaly detection, churn and predictive modeling, A/B testing.
FastAPI services, event-driven microservices, CI/CD, and observability.
Saint Peter’s University
University of Mumbai
A production LLM/RAG pipeline I design and operate - from raw data to a served, monitored endpoint that stays fast and cheap under load. Pick a stage.
Event-driven intake that survives the real world: external APIs behind rate limits, retries with backoff, schema normalisation. ~500K events a day arrive without anyone getting paged.
A transformer alternative built on hierarchical latent state-space recurrence - trading quadratic attention for recurrence that holds long context on hardware that fits in a pocket.
Presented this line of work on stage in 2026 - see below.
Grouped the way the system is layered, not alphabetically.
AWSGoogleMicrosoftNVIDIATensorFlowIBMPyTorchStanford
Ten of them, all public - nine on GitHub and one on the VS Code Marketplace.
AI harness for VS Code
An AI harness living in the IDE - it plans before it codes, reads the codebase and follows call sites, runs commands and reacts to failures, and presents every change as a diff you approve. Extends through MCP. Live on the VS Code Marketplace.
Offline LLM assistant for Android
On-device inference with zero cloud dependency. Runs LLaMA 3.1, IBM Granite, and Mistral via 4-bit GGUF at ~8 tok/s, with on-device RAG over your own documents.
Local-first context compression
Context compression for LLM applications that runs locally - 50–90% fewer tokens for the same answers, with nothing shipped to a third-party service to get there.
Long-context LLMs on edge
Reference code for my TechRxiv preprint: hierarchical latent state-space recurrence that democratises long-context LLMs on edge devices, as a transformer alternative.
Production RAG on Kubernetes
Demo from my KCD New York 2026 talk - deploying, autoscaling, and observing production RAG pipelines on Kubernetes with vLLM and vector search.
Real-time AI pipelines
Embedding LLMs and RAG directly into Apache Beam transforms for low-latency inference on high-velocity data streams.
Open-source outbound agent
An open-source AI agent for outbound leads - the research, qualification, and follow-up loop handled by a tool-using agent rather than a sequence of manual steps.
Open-source AI note-taking
An open-source AI note-taking app in the Granola mould, written in Rust - meeting notes captured and structured by a model, without the notes leaving a proprietary product.
Agentic personal finance
An agentic-AI finance tracker that ingests transactions and reasons over them with tool-using agents to surface spend insights.
AI-or-human text detection
Detects whether a piece of text is AI-generated or human-written using a perplexity-based metric, served through a lightweight web app.
Numbers from systems that shipped and stayed up. Every one of them came with a dashboard someone else had to trust.
multi-agent + retries
Kafka pipeline
FastAPI on Kubernetes
NLP ticket routing
predictive ML
A/B testing programme
In 2026 I took the hard part of applied AI to two New York stages - not the demo, but making it run reliably, cheaply, and at scale.
How retrieval-augmented generation earns its uptime: autoscaling vLLM, vector search, and observability on Kubernetes so a RAG system stays fast and affordable when traffic stops being a benchmark and starts being real.
Bringing LLM inference and RAG inside the pipeline itself - embedding models directly into Apache Beam transforms so streaming data can reason in real time, delivering low-latency intelligence on high-velocity data without ever leaving the stream.
I will be speaking at Blockchain Week - UNGA Edition 2026, the ten-day independent industry gathering held in New York alongside the UN General Assembly (September 10–19), with tracks spanning Bitcoin, AI agents, energy, and the space economy. The conference will announce the session title and slot; the talk sits where my work does - production AI systems, and what it takes to run them reliably once the demo is over.
A mentor, a Microsoft engineer, and a reader, in their own words.
Samir is an inventor. His work with agentic AI, open-source contributions on Hugging Face, and a locally-run IBM Granite LLM app for Android set him apart. There were times you’d be hard-pressed to tell who was the professor.
Our team had spent weeks trying to structure a complex agentic-AI workflow. Samir stepped in, grasped the entire system, and had it fully running in under a week.
LOVED your productivity paradox feature. Just read your Medium article and I loved it - particularly the guilt piece! I talk a lot about maximising our human capacity to thrive in this era, and that was a new one.
Open to AI/ML Engineering, Data Science, and Python roles, plus research collaborations and consulting. Based in New York, shipping worldwide.
Fastest route - straight to my inbox