A few frames from along the way
From an award ceremony back home to the simulations running in my research right now.
Building AI systems that see, reason, and act.
I work across computer vision, generative AI, and multi-agent systems — from vision-language models that explain a satellite's components in orbit, to language models that command a drone swarm on the ground.
"I build systems that don't just output an answer — they show their work."
I'm an M.Tech student in Systems & Control Engineering at IIT Bombay, working at the intersection of computer vision, generative AI, and multi-agent systems. My thesis teaches machines to look at a satellite's components and explain what they see — pairing detection models with language models so the reasoning stays visible, not just the output.
Before IIT Bombay, I graduated top of my cohort in Computer Science & Engineering at Katihar Engineering College. Since then I've shipped a production RAG pipeline for a hiring startup, trained a language model to command a drone swarm, and co-authored work that helped bring a machine-learning framework into cancer research.
Three things worth pausing on
Money that backed an idea, a paper that reached a real journal, and data that's now in other people's hands.
Marks that took real work to earn
Medals, grades, and rankings collected across six years of coursework and competitive exams.
Leadership & teaching
Responsibilities I've taken on alongside research — from running placement operations for the institute to helping others learn generative AI.
From seeing shapes to explaining what it sees
Two connected chapters of my M.Tech — a seminar in geometric vision, and a thesis teaching a vision-language model to reason about what it detects in plain language.
- Extend the framework across additional visual domains beyond the current satellite testbed.
- Integrate retrieval-augmented LLM reasoning for grounded interpretation, anomaly analysis, and natural-language reporting.
Where I've put this to work
Helped turn a firehose of job postings into something searchable, explainable, and fast enough to run in production.
- Processed 10,000+ job postings through an end-to-end ML job-intelligence pipeline with automated ETL.
- Evaluated post-trained semantic classification on 3,000+ labeled samples generated by Qwen 3-8B-Instruct.
- Built a hybrid RAG system with PostgreSQL, ChromaDB, Sentence Transformers, and LangChain reasoning.
- Productionized the pipeline with FastAPI and Docker, powering a live dashboard for AI-driven job search, analytics, and replies.
Built a language model that can command a drone swarm, and benchmarked it against the alternatives.
- Developed a hybrid LLM-PPO multi-agent commander over 10,000+ scenarios, pseudo-labeled using GPT-OSS-120B and Llama-3.3-70B-Versatile.
- LoRA fine-tuned Qwen2.5-0.5B-Instruct to 91.2% token accuracy.
- Benchmarked 7 architectures across LLMs, swarm intelligence, and RL: 67.9% action accuracy versus 60.7% for LLM-only and 57.1% for PSO, with 0.986 PPO explained variance.
- Recorded the academic video for CSDS119, Computer Vision and Image Processing, at SVNIT Surat.
Selected projects
Filter by focus area to see how each project fits together.
A visual QnA system that reasons like a person looking at a scene: YOLO11x and DPT-Small handle grounding and depth, while Llama-3.3-70B and Llama-4-Scout-17B decide which of 10 specialized tools to call next. Reached ~7–9s end-to-end latency at 0.35 confidence and 0.60 IoU, deployed on Streamlit.
A RAG assistant built on n8n with a Google Drive-synced knowledge base that ingests updates automatically, served through a ChatGPT-style interface with text and voice. Averages ~1.42s end-to-end, split between ~0.15s retrieval and ~1.23s generation.
A DNS caching simulator with LRU and TTL policies under configurable network conditions across 5,000+ requests, tracking hit ratio, P95 latency, upstream traffic, and retransmissions.
Simulated annealing for TSP, minimizing tour cost through stochastic search, with route and convergence visualizations to make the optimization process legible on a 40-city instance.
A CVAE for class-guided MNIST synthesis across all 10 digits, trained on a 2D latent space with reconstruction loss and KL divergence, with decoder-only inference deployed on Render for real-time generation across a [-3, 3]² latent grid.
A life-expectancy prediction pipeline with a 70:15:15 split and leakage-safe preprocessing, evaluated across 7 models under 5-fold cross-validation. Extra Trees came out ahead at 0.97 test R² and 1.66 RMSE, with 441 predictions wired into Power BI for ongoing evaluation.
A 3-stage speech pipeline chaining ASR, classification, and sentiment analysis with Whisper-Small and RoBERTa, reaching 96.8% classification accuracy and 92.6% sentiment confidence, running entirely in Streamlit with zero external inference APIs.
A DataMatrix-based tablet traceability system for instant medicine identification and expiry tracking, funded with a ₹4.5 lakh government grant after scoring 95.75 out of 100 in the NAIN 2.0 program.
Written & filed work
The stack behind the work
Let's talk about vision, language, or systems.
Open to research collaborations, internships, and conversations about computer vision, generative AI, and multi-agent systems.


