Chandan Kumar.AI/ML Engineer
I build production AI systems across large language models, retrieval pipelines, and computer-vision tooling that move from notebooks to real users. Currently focused on LLM fine-tuning, RAG, and multimodal perception.
Things I've shipped recently.
A selection of AI tools: open-source libraries, research interfaces, and shipped products.
A Python library and MCP server that detects and masks PII before it reaches LLMs. Hybrid detection engine combines regex with Microsoft Presidio and spaCy NER, using reversible tokenization and a per-request in-memory vault. Ships as a library, FastAPI middleware, CLI, and MCP server for Claude Code.
BitWiser-1q1
One-click 1-bit (IQ1_S) compressor for any local LLM on Apple Silicon. Auto-discovers Ollama and on-disk GGUF models, runs importance-matrix calibration, and quantizes via llama.cpp with Metal — shrinking Llama-3.1-8B from 16.1 GB to 2.19 GB (7.4×) at ~51 tok/s on 3.1 GB of RAM. Side-by-side playground streams live tokens/sec, RAM, and bandwidth.
Facematch
Face verification pipeline that compares faces across images, PDFs, and Excel documents. RetinaFace detection with quality filtering and automatic rotation handling, ArcFace embeddings via DeepFace, and cosine-similarity matching with configurable thresholds — producing detailed reports with per-document confidence scores.
Vocat
Real-time voice interview agent. Google-Meet-style UI with WebRTC audio, Whisper STT, GPT-4o reasoning, and ElevenLabs TTS. Streaming responses with VAD-based turn-taking for natural conversational latency.
Whisper Hindi ASR
Fine-tunes OpenAI Whisper-small for Hindi speech recognition using LoRA (PEFT) on Mozilla Common Voice 17, training under 2% of parameters to measure WER/CER gains over the baseline. Type-hinted library with YAML-driven config, runs as fast CPU/MPS smoke tests locally and full fp16 training on a Colab T4.
I bridge the gap between research and production.
Currently a final-year CSE student at VIT Vellore, with internships at ISRO, IIT Ropar, IIT Bombay, and AI startups. I obsess over the path from notebook to deployed system: latency, model size, retrieval quality, and the boring infra glue that makes models actually useful.
My focus is the part most teams skip: the last mile. Quantizing a 14GB model to 4GB without losing accuracy. Cutting inference under 200ms. Replacing a brittle OCR script with a YOLO + PaddleOCR pipeline that handles real-world engineering drawings.
Where I've been training & shipping.
NeuroFin.ai
Onsite · Bengaluru- •Building an intelligent KYC verification system using computer vision and deep learning to automate document validation and identity matching.
- •Designed a robust pipeline that classifies document types, extracts key fields, and performs face matching for secure user onboarding.
IIT Bombay · TuroCrate.ai
Remote · Mumbai- •Designed and shipped ML systems at TuroCrate, focused on production data pipelines and inference services.
- •Owned the full model lifecycle: training, evaluation, deployment, and observability across the platform.
IIT Ropar · Annam.ai
Onsite · Punjab- •Built a CNN plant-disease classifier reaching 93% accuracy via ResNet-50 transfer learning on 87K+ images.
- •Developed an ML recommendation engine and integrated NLP-based sentiment analysis for agricultural news.
ISRO · LPSC
Onsite · Kerala- •Created a custom OCR pipeline combining YOLOv5 and PaddleOCR to digitize 500+ engineering drawings.
- •Applied image preprocessing techniques that improved model robustness by 15%.
Let's build something remarkable.
Open to AI/ML roles, interesting research collaborations, and ambitious shipping projects. Reach out, I read everything.
cml.codes@gmail.com→