background
Education
Master of Science
Carnegie Mellon University, School of Computer Science
Artificial Intelligence and Innovation · GPA: 4.08 / 4.0
Coursework
Introduction to Machine Learning (PhD)
Advanced Deep Learning (PhD)
Computer Vision
Advanced Natural Language Processing (PhD)
AI Engineering
Learning for 3D Vision
Deep Learning Systems
Generative AI
Computer Systems
Bachelor of Technology
Delhi Technological University
Electronics and Communications Engineering · GPA: 9.14 / 10.0
Coursework
Deep Learning
Computer Vision
Data Structures
Algorithms
Object-Oriented Programming
Database Management Systems
Experience
ML Intern, Neuron Scalable Training
- Enabled diffusion and flow-matching model training on AWS Trainium by onboarding SOTA models (Stable Diffusion, FLUX, DiT, Wan) onto Neuron, Annapurna's native PyTorch backend.
- Implemented and optimized PyTorch ATen operators for the Neuron backend, extending native operator coverage and efficiency.
- Validated long-horizon training dynamics between Neuron and CUDA across text-to-image, text-to-video, and diffusion language models, ensuring convergence, numerical stability, and quality parity.
AI Entrepreneur and Researcher
- Architected a generative AI pipeline leveraging diffusion transformers and cross-modal alignment to synthesize photorealistic, audio-driven video avatars with high temporal coherence and lip-sync accuracy.
- Led a 5-member ML engineering team in developing an MVP launched on Product Hunt, implementing optimization strategies for real-time avatar generation and conducting iterative evaluations to improve perceptual realism and latency.
- Formulated go-to-market strategy via business plans, pitch decks, and financial models; showcased at VivaTech (Paris) and Bits & Pretzels (Munich) to attract partners and investors.
Computer Vision Engineer
- Engineered one-shot anomaly detection pipelines leveraging feature embedding similarity and memory-bank architectures, achieving over 99% detection accuracy with under 1% false positives in real-time defect identification on manufacturing lines.
- Led design and integration of a transformer-based OCR module into the anomaly detection stack, replacing legacy classical CV approaches; delivered near-perfect text recognition accuracy for reliable automated inspection in production.
- Performed remote rollout and on-device validation of CV models on Linux-based embedded hardware, ensuring robust edge inference.
Undergraduate Research Intern
- Explored audio-to-video facial synthesis and fingerprint enhancement using generative AI; submitted 3 research papers to journals and conferences, later pursuing incubation to translate research into products.
- Adapted the Latent Image Animator model for audio-conditioned video synthesis via feature mapping of audio encodings to motion space; transitioned to diffusion-based models for improved temporal coherence and realism.
- Developed novel self-supervised losses and integrated FFT-based convolutions in FG-GAN to enhance fingerprint image quality for biometric recognition.
Skills
AI & ML Frameworks: PyTorch, CUDA, JAX, TensorFlow, Keras, HuggingFace (Transformers, Diffusers, PEFT), Weights & Biases
Generative AI: Diffusion Models (DDPM, Flow Matching, CFG), Stable Diffusion, FLUX, DiT, Wan2.1, LLMs (fine-tuning, LoRA, RLHF, RAG), GANs
Computer Vision: CLIP, ViT, DINO, Object Detection, Semantic Segmentation, Image Enhancement, OpenCV
ML Systems: AWS Trainium, Neuron SDK, Distributed Training (DDP, FSDP, ZeRO), Mixed Precision (BF16/FP8), HPCs, Docker
Languages & Tools: Python, C/C++, CUDA, Bash, SQL, Git, Linux, Conda, Jupyter