machine learning engineer bengaluru, in

Dhriman Deka©26

Mostly trying to make small things work a little better than yesterday.

llm evaluation · retrieval ·
and the mlops around them

↓ scroll, slowly

01 — About

I build and evaluate LLM systems — retrieval pipelines, eval harnesses, and the MLOps around them. Most weeks: shipping models to production, then measuring what actually breaks.

Working on
LLM evaluation, retrieval
Stack
Python, PyTorch, vLLM
Reading
Anthropic interp papers
Listening
Caribou — Suddenly

02 — Work

01 / 03

Data Science & ML Engineer

Probe42 — Bengaluru

Productionising ML — moving models from Jupyter notebooks onto EC2 as reliable services with monitoring, retraining, and clean interfaces.

Built a BERT-based semantic search system and a generative model for synthetic compliance data, now used across internal tooling. Own the full path: data pipelines, training, evaluation, deployment, and the boring glue in between.

2024 — now
02 / 03

Research Intern, LaRGo LLM Group

with Prof. Kripa Bandhu Ghosh — Kolkata

Fine-tuned a multilingual hate-speech classifier across several Indian languages on ~1.5M tweets.

Learned how much of applied LLM research is careful eval design, dataset curation, and reproducible pipelines.

Feb — Aug 2024
03 / 03

B.S. Economics

IISER Bhopal

Econometrics and statistics give a strong mathematical foundation for the ML work — especially around causal inference, portfolio optimisation, and reading models honestly.

A habit, from economics, of not trusting clean-looking numbers.

03 — Projects

Selected
projects

Five things, built end to end — from data to deploy.

scroll →

202501 / 05

Meridian

A clinical-grade mental health agent. LangGraph orchestrator with deterministic safety gates, MentalRoBERTa ONNX screening, and session-scoped ChromaDB RAG — ephemeral by design.

  • LangGraph
  • MentalRoBERTa
  • ONNX
  • ChromaDB
  • FastAPI
202402 / 05

BERT search for an internal tool

Replaced a clunky filter UI with semantic search. Helped the team find things faster; helped me find a lot of bad data.

  • BERT
  • PyTorch
  • FastAPI
  • Docker
202403 / 05

Multilingual hate-speech detector

Fine-tuned on 1.5M tweets across a few Indian languages. Decent F1, very humbling failure cases.

  • Transformers
  • Hugging Face
  • PyTorch
202304 / 05

Pension portfolio optimiser

Monte Carlo + classical optimisation on Indian pension data. A chance to think carefully about what risk actually means.

  • Python
  • Pandas
  • Scikit-learn
202305 / 05

Synthetic PAN/GST generator

A small generative model for compliance test data. Useful, ugly, shipped.

  • Python
  • ML

04 — Stack

A small, boring stack —
chosen mostly because it gets out of the way.

Python/ PyTorch/ Hugging Face/ vLLM/ FastAPI/ Transformers/
RAG/ LoRA · QLoRA/ Pandas/ SQL/ Docker/ AWS/
TypeScript/ JAX/ Triton/ MLflow/ Airflow/ Spark/

05 — Writing

Notes I’ve written down — partly to remember, partly to be corrected.

2026

Invited reviewer — SD4H Workshop @ ICML 2026

Nominated by the program chair committee to review emerging ideas in sustainable development for health.

openreview ↗
2024

A comparative look at NPS and UPS from an employee’s perspective

A small paper on Indian pension reform. Mostly accounting and assumptions.

ssrn ↗
2024

Implementing a DQN agent in Ludo: an exploration of reward structures

10,000 episodes, ~9k lines of code, 43% win rate. Notes on shaping rewards for a game full of luck.

medium ↗
2024

Building a transformer for stock prediction: challenges, overfitting, and why it lost money

53.7% directional accuracy, and an honest write-up of where the model breaks and why markets don’t forgive it.

medium ↗
Schwarzschild · null geodesics