Research themes
What I focus onTransformer compilers, GPU-accelerated inference, adversarial LLM evaluation, explainable AI, time-series modeling, and applied machine learning systems that reveal where models actually break.
I build systems that ▍
I build AI systems that pair research rigor with deployable engineering. My recent work spans transformer compilers, GPU-accelerated LLM inference systems, adversarial evaluation infrastructure, and trustworthy AI platforms, with hands-on work across CUDA, Triton, PyTorch, FastAPI, and full-stack ML systems.
Transformer compilers, GPU-accelerated inference, adversarial LLM evaluation, explainable AI, time-series modeling, and applied machine learning systems that reveal where models actually break.
Python, C++, CUDA, Triton, and TypeScript across ML pipelines, compiler/runtime systems, backend services, experiment tracking, and cloud infrastructure.
Python, Java, C/C++, SQL, JavaScript, HTML/CSS, R
PyTorch, TensorFlow, Keras, Transformers, scikit-learn, NumPy, Pandas, Matplotlib
CUDA, Triton, C++20, pybind11, custom IRs, operator fusion, kernel selection, memory/execution planning, paged KV caching, quantization
LangChain, ChromaDB, RAG, vector databases, PyPDF
Node.js, FastAPI, Flask, REST APIs, Docker, Kubernetes, Linux, Git, CI/CD, Nginx
AWS EC2, S3, IAM, Lambda, SageMaker, CloudWatch, PostgreSQL, Redis, ETL, schema design
TDD, API integration, DSA, system design, performance optimization, code reviews, debugging
Built ReliaGuard Studio and BreakPoint for human-AI reliance detection and adversarial LLM evaluation.
Researched multimodal toxicity mitigation and anomaly detection with reproducible experiment workflows.
Built an end-to-end LSTM satellite clock-bias prediction platform for more reliable navigation timing forecasts.
Developed Node.js-backed LLM pipelines for chat, knowledge graphs, and retrieval-augmented medical QA.
Master of Science in Computer Science and Engineering - Santa Cruz, California
Bachelor of Technology in Computer Engineering - Ahmedabad, India
Inspectable compiler and runtime for GPU-accelerated decoder-only transformer inference: imports Llama, Mistral, and Qwen-style models into a custom IR, applies transformer-specific optimizations, plans memory and execution, selects CUDA kernels, and runs through a reference runtime with vLLM-compatible serving concepts.
Failure-to-eval compiler for LLM systems that mines real failure modes into versioned adversarial regression suites with DSL-defined traps and rubrics, multi-judge calibration, RAG/tool simulators, CI gates, and FastAPI/Next.js dashboards.
Next.js, TypeScript, and FastAPI observability platform with SDKs, guardrail APIs, review queues, and conformal neuro-symbolic models to detect and audit overreliance and underreliance in human-AI workflows.
Universal self-improving AI research agent that discovers papers, generates hypotheses, runs experiments in a sandbox Python environment, and updates its reasoning strategy using long-term graph memory.
Real-time AI meeting copilot desktop app that captures system audio, transcribes conversations, and streams context-aware responses using the OpenAI API with screenshot and document context support.
Benchmarked TimeGPT, Chronos, and Time-MOE across multiple datasets, showing where traditional statistical and deep learning models can still outperform TSFMs.
Automated pipeline for generating structured knowledge graphs from unstructured text using large language models, with evaluation across GPT-4, LLaMA-2, and BERT.
Retrieval-augmented medical assistant built to provide accurate and reliable answers grounded in medical PDFs.
LSTM-based clock-bias forecasting with a practical preprocessing pipeline for navigation reliability.
Structured knowledge graph generation from unstructured text, evaluated for semantic quality and GraphRAG readiness.
Retrieval-augmented medical assistant for grounded answers from medical literature PDFs.
Evaluation of time series foundation models for anomaly detection and prediction under realistic benchmarking settings.
For research collaborations, internships, or ML engineering conversations, send a short note with context and one link.