š Systems Projects & Technical Deep-Dives
Showing all 11 projects
Bare-Metal KVM Hypervisor & Virtualization Engine
Course Project | CS 695: Topics in Virtualization and Cloud Computing | Prof. Purushottam Kulkarni
- Engineered a userspace hypervisor using Linux KVM API and mmap to manage isolated guest vCPU contexts.
- Implemented guest-to-host virtual address translation by handling KVM I/O-port exit traps in userspace.
- Orchestrated time-sliced vCPU scheduling using a circular buffer to synchronize Producer/Consumer VMs.
Conductor: Docker-Lite Container Runtime & Image Engine
Course Project | CS 695: Topics in Virtualization and Cloud Computing | Prof. Purushottam Kulkarni
- Engineered a container runtime using namespaces and cgroups v2, achieving 73x faster startup than Docker.
- Accelerated image builds by 85x via SHA-256 content-addressable storage and OverlayFS layer caching.
- Implemented container networking using veth pairs and iptables routing for a C/Flask microservice.
Linux Kernel Modules: MMU Walks, Ioctl Drivers & Telemetry
Course Project | CS 695: Topics in Virtualization and Cloud Computing | Prof. Purushottam Kulkarni
- Implemented 5-level page walks for VA-to-PA translation and 2 MiB Transparent Huge Page (THP) accounting.
- Developed character drivers using ioctl interfaces for physical-memory writes and process tree reparenting.
- Exposed page-fault & memory statistics via procfs/sysfs, comparing O(1) RSS accounting with O(N) page walks.
RustOS: Bare-Metal Preemptive Multitasking Microkernel
Course Project | CS 744: Design and Engineering of Computing Systems | Prof. Purushottam Kulkarni | Rust, x86_64
- Built a 64-bit x86_64 Rust kernel, configuring the GDT and IDT for interrupt and exception handling.
- Implemented physical frame allocation and heap allocation using bootloader-provided memory mappings.
- Orchestrated Round-Robin multitasking via PCBs, context switching, and Mutex/Semaphore synchronization.
- Developed a VGA driver, an interactive shell, and a hierarchical in-memory VFS with CRUD support.
MyRedis: In-Memory Key-Value Database
Self Project | C++, Linux Sockets, Multithreading, I/O Multiplexing
- Engineered a non-blocking event-driven Redis-like server using multiplexed I/O and TCP pipelining.
- Designed an incremental-rehashing key-value store and dual-indexed sorted sets for O(log N) range queries.
- Implemented active TTL key eviction, LRU connection pruning, and asynchronous memory cleanup.
Workload-Aware Index Selection for PostgreSQL
M.Tech Thesis | Guides: Prof. Suraj Shetiya & Prof. Om Damani
- Developed a custom PostgreSQL extension using Executor Hooks and shared-memory hash tables to collect column-level UPDATE statistics (<2% runtime overhead across 50K concurrent DML queries).
- Designed a B-tree write-penalty model using planner costs and UPDATE statistics for index update overhead.
- Implemented the ICDE'19 recursive multi-attribute index selection algorithm in PostgreSQL via HypoPG.
- Investigated workload-aware database optimization and self-driving DBMS adapting to workload forecasts.
Attention Mechanisms & LLM Inference Optimization
Course Project | CS 794: Systems for Machine Learning | Prof. Mythili Vutukuru
- Accelerated the decode phase by 29.1x GPU and 9.8x CPU using KV caching to avoid redundant computation.
- Designed a radix-tree shared-prefix KV cache to eliminate redundant prefill, gaining a 1.25x GPU speedup.
- Developed FlashAttention in CUDA with online-softmax tiling, scaling VRAM footprint from O(N²) to O(N).
CRISP: Critical Slice Prefetching for OoO Processors
Course Project | CS 683: Advanced Architecture | Prof. Biswabandan Panda
- Architected critical-slice prefetching in ChampSim to identify and prioritize latency-critical dependency chains.
- Redesigned the OoO scheduler with priority issue queues to issue critical slices while preserving ROB ordering.
- Improved Limit Order Book matching engine IPC by 4.98% evaluated on Intel PIN-generated traces.
GPU-Optimized Matrix Multiplication & MLP Forward Pass
Course Project | CS 794: Systems for Machine Learning | Prof. Mythili Vutukuru
- Profiled MLP forward-pass performance across SIMD, cache tiling, prefetching, and CUDA shared memory.
- Applied cache tiling (26.1x) and AVX2 SIMD (55.8x) to CPU matmul, reducing L1/L2 cache misses by 43x.
- Deployed shared-memory tiling in CUDA matmul (1.7x speedup), boosting GPU utilization from 13.6% to 64.5%.
RAG Documentation Assistant: Hybrid Search & Reranking Pipeline
Self Project | Python, PostgreSQL, pgvector, Gemini API
- Built an ingestion pipeline chunking technical docs into PostgreSQL using pgvector (HNSW) and GIN indexes.
- Designed hybrid retrieval combining vector search and full-text search with cross-encoder reranking.
- Generated citation-grounded LLM responses with source references to reduce hallucinations.
Sherlock: Local-First Codebase RAG & Runtime Debugger
Self Project | Python, Tree-sitter, LanceDB, FastEmbed, Ollama
- Engineered AST-aware code chunking across 11 languages with Tree-sitter, isolating function and class boundaries.
- Built in-process vector retrieval using embedded LanceDB with FastEmbed ONNX embeddings and SHA-1 incremental indexing.
- Developed three-way rule-based query routing, stack trace interval matching, and offline Ollama LLM diagnosis.
š ļø Technical Skills
C++ / C
Rust
Python
CUDA
Bash / Shell Scripting
SQL & PostgreSQL
Linux Kernel & MMU
KVM Hypervisor
Namespaces & cgroups v2
OverlayFS
iptables / Netfilter
x86_64 Paging
AWS
Docker
Kubernetes
Ansible
Linux perf & GDB
Git