Core Specialisations
Large Language Models
Architecture internals, reasoning diagnostics, failure taxonomy, RLHF/GRPO fine-tuning, prompt engineering at scale.
Generative AI
Text generation pipelines, instruction-tuned models, retrieval-augmented generation, agentic LLM workflows.
Deep Learning
Transformer architectures, attention mechanisms, LoRA/QLoRA adaptation, model evaluation and benchmarking.
Reinforcement Learning
Policy optimisation, reward shaping, GRPO/PPO for language models, TRL library, multi-step reasoning incentivisation.
Indic NLP · BharatGen
Low-resource multilingual benchmarking, cross-lingual reasoning evaluation, native fluency in Hindi and Telugu.
Mathematical AI
Linear algebra, probability theory, statistical inference, optimisation, calculus — applied to ML model design.
Keyword Index
BharatGenLarge Language ModelsGenerative AILLM Reasoning & Failure AnalysisGRPOReward ShapingDeep LearningTransformer ArchitecturesRetrieval-Augmented GenerationReinforcement LearningIndic NLPChain-of-Thought PromptingInstruction Fine-TuningGraph Neural NetworksLinear AlgebraProbability TheoryCalculus & OptimisationStatistical InferencePyTorchHuggingFaceTRLPythonC++CUDAParam Cluster
Research
M.Tech Thesis — Multi-Teacher On-Policy Distillation for Verifiable and Unverifiable Domains
RL-trained domain experts, shared SFT checkpoint, per-prompt routing, token-level reverse-KL distillation, exposure-bias elimination, multi-round iterative distillation, verifiable + unverifiable domain unification
R&D Report — Empirical Study of LLM Reasoning Failures on English Benchmarks
4-model evaluation, 199-question benchmark, 10-category failure taxonomy, premise-order sensitivity, multi-step error propagation, GRPO reward shaping
Seminar — Reasoning in LLMs
Chain-of-Thought, Tree-of-Thoughts, recursive decomposition, RL-based reasoning, inference-time vs. training-time trade-offs, exploration-efficiency analysis
Projects
Multimodal GraphRAG for Long-Document QA
VLM captioning, knowledge graph construction, entity resolution, 2-hop traversal, community detection, MMLongBench-Doc
Reasoning Failure Diagnosis in Indic-Language LLMs
16-category failure taxonomy, 17B Indic LLM, MGSM, Hindi/Telugu/Bengali
Tapestry: AI Alliance
Federated fine-tuning, Llama 3.1 8B, multi-GPU multi-node training, WVS + IVM evaluation pipeline
Agentic Research Assistant with Multi-Source Synthesis
LangGraph state machine, autonomous agent nodes, citation-backed synthesis
MNIST Compression Robustness Analysis
Pruning, distillation, INT8 quantization, LeNet-5
Paged KV-Cache Scheduler for LLM Inference
Continuous batching, paged KV cache, GPU scheduling
CivicEase – Civic Services Aggregation Platform
Flask, SQLAlchemy, Bcrypt authentication, RSS ingestion pipeline
Multi-Threaded Key-Value Store with LRU Cache & Load Benchmarking
C++, PostgreSQL, LRU cache, mutex-guarded connection pooling
Skills & Technologies
LLMs & Gen AI
Reasoning analysis & benchmarking, instruction fine-tuning (LoRA/QLoRA), GRPO/PPO reward shaping, retrieval-augmented generation, Chain/Tree/Buffer of Thought, prompt engineering & evaluation
Deep Learning
Transformer & attention mechanisms, sequence-to-sequence models, graph neural networks, convolutional & recurrent nets, model compression & quantisation, transfer learning
ML & RL
Supervised & unsupervised learning, policy-gradient reinforcement learning, Bayesian methods, ensemble & boosting methods, evaluation metrics, experiment tracking with Weights & Biases
Mathematics
Linear algebra & matrix decompositions, probability theory & stochastic processes, statistical inference & hypothesis testing, multivariate calculus & gradients, convex optimisation, information theory
Frameworks & Libraries
PyTorch, TensorFlow, HuggingFace Transformers, TRL, PEFT, LangChain, LlamaIndex, scikit-learn, NumPy, Pandas, NLTK, spaCy, IndicNLP, MLflow, Verl, PrimeRL, Slakshana
Programming & Systems
Python, C, C++, Java, JavaScript, HTML, CSS, SQL, Bash, CUDA/GPU programming, Linux, Git, Docker, Kafka, PostgreSQL, MySQL, LaTeX, Param Supercomputer, OS, Networks, COA, DBMS, DSA
Education
M.Tech — Computer Science & Engineering
MTP under Prof. Ganesh Ramakrishnan · BharatGen Project · Param Compute Cluster Access · specialisation in LLMs, Generative AI, and Indic NLP
B.Tech — Computer Science & Engineering
Strong foundation in algorithms, mathematics, and systems programming · 9.02/10 CGPA
Higher Secondary — Science (PCM)
~98%
Secondary Education (SSC)
10/10 CGPA
Grounded In
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Wei et al., 2022
Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Yao et al., 2023
Buffer of Thoughts: Thought-Augmented Reasoning with LLMs, Yang et al., 2024
Algorithm of Thoughts: Enhancing LLM Exploration via Algorithmic Reasoning, Sel et al., 2023
Multi-Step Action Planning for Complex Reasoning in LLMs
DeepSeek-R1: Incentivising Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek AI, 2025
Premise Order Matters in Question Answering for Large Language Models, Chen et al., 2024
On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes, Agarwal et al., 2023
MiniLLM: Knowledge Distillation of Large Language Models, Gu et al., 2023
Reinforced Multi-Teacher Selection for Knowledge Distillation, Yuan et al., AAAI 2021
Exploring Knowledge Purification in Multi-Teacher Knowledge Distillation for LLMs, Jin et al., ICLR 2026
Achievements
GATE Qualified
GATE CS 2025 — All India Rank 289 (~99.99 percentile), cleared in B.Tech 4th year. GATE CS 2024 — All India Rank 606, cleared in 3rd year.
Researcher, BharatGen
Selected to contribute to IIT Bombay's national flagship Indic-language AI initiative, under Prof. Ganesh Ramakrishnan, with AWS cluster GPU access.
Best Student Award
Top-1 finish — Code Craft
CodeWiz, organized by ACM.
Participated in ICPC
International Collegiate Programming Contest — an international coding olympiad.
Positions of Responsibility
Reasoning Team Member — Machine Learning Lab, IIT Bombay
Fresher guidance, core ML fundamentals, LLM training phases, lab research showcase
Teaching Assistant — CS 769: Optimization in Machine Learning
Research-paper-based projects, skill-based interviews, exam invigilation
Teaching Assistant — CS 207: Discrete Structures
Doubt-clearing sessions, assignment grading, evaluation support, Moodle management
Teaching Assistant — CS 228: Logic for Computer Science
Exam leadership, grading, invigilation, 200+ student mentoring
Student Companion — Institute Student Companion Programme (ISCP)
Rigorous selection process, fresher support, academic & socio-cultural guidance, holistic well-being
Academic Coursework
M.Tech — IIT Bombay
Advanced Machine Learning, Deep Learning & Neural Networks, Natural Language Processing, Foundations of Reinforcement Learning, Graph Learning & Network Analysis, Advanced Algorithms & Complexity, Computer Vision, Research Methodology
B.Tech Foundation
Data Structures & Algorithms, Object-Oriented Programming, Database Management Systems, Computer Networks & Security, Operating Systems, Software Engineering, Theory of Computation, Probability & Statistics