AI Engineer — Multi-Agent Systems

SathwikMarupaka

I build multi-agent AI systems that ship, not just chat.

Portrait of Sathwik Marupaka

Trajectory

2020

B.Tech, IIT Dhanbad

  • Electronics & Communications Engineering
  • Advanced Algorithms · Graph Algorithms · Principles of AI · Machine Learning & AI · Probability & Statistics

2022 · mentee

Microsoft Engage

  • Selected as mentee, summer 2022

2023 · SDE intern

Technophilia

  • SDE intern, summer 2023 — see Experience for details

2024–2026

MS CS, GWU

  • Machine Intelligence & Cognition track
  • Advanced Machine Learning · Neural Networks & Deep Learning · Artificial Intelligence · Data Mining · Cloud Computing

2025 · AI Engineer

DSSD, GWU

  • AI Engineer, Jan–Dec 2025 — see Experience for details

500+ problems, C++

Algorithms & DSA

  • 500+ problems on LeetCode and GeeksforGeeks
  • Data structures and algorithms in C++

4 podium finishes

Hackathons

  • Agents for Impact (NVIDIA), Aug 2025 — top 5
  • HackWithDC at AWS, Feb 2026 — top 3
  • ProductTank Lovable Hackathon, Feb 2026 — top 3
  • DevFest DC Build-a-thon, Aug 2026 — 2nd place

Apr 2026

GeorgeHacks

  • GWU School of Engineering & Applied Science, Apr 2026
  • Produced NutriRX

CRE · Compliance · NutriRX

Flagship projects

  • Multi-Agent CRE Underwriting · AI Compliance Auditor · NutriRX
  • All three are live and runnable in the Demos section above

multi-agent systems

Now

Experience

AI Engineer

Data Science for Sustainable Development (DSSD), GWU

Jan 2025 – Dec 2025

PythonCatBoostNetworkXFAISSDBSCAN
  • Built and trained an hourly demand-forecasting model for Capital Bikeshare, processing 4M+ 2023 trip records across four DMV counties joined with station geospatial data and EPA hourly temperature readings
  • Achieved 78.15% accuracy within ±1 trip using CatBoost on a time-aware split (Jan–Oct train, Nov–Dec test); segmented 182 high-usage stations into four operational hotspots via DBSCAN to prioritize redistribution routes for nonprofit partners
  • Diagnosed and removed a leaking historical-average feature that was suppressing model learning, then ran a three-way ablation isolating temperature's contribution at +4 percentage points (74% → 78%)
  • Led internal research benchmarking knowledge-graph against vector retrieval for LLM code generation — an 855K-node / 1.96M-edge AST-derived graph (NetworkX) against a FAISS index over 455K Python files. Raised Pass@1 on LiveCodeBench from a 64.5% no-retrieval baseline to 84.25%, and isolated why graph traversal underperformed at 77.50%: keyword-dependent lookup fails on natural-language problem statements

Software Developer Intern

Technophilia Solutions

May 2023 – Jun 2023

PythonPyTorchHugging FaceMamba
  • Engineered an evaluation pipeline comparing classical models (SVM, Logistic Regression), a voting-based hybrid ensemble, and Transformers (DistilBERT, XLM-RoBERTa); XLM-RoBERTa reached 63.70% accuracy on bilingual sentiment tasks
  • Integrated the Mamba state-space architecture to sidestep the quadratic attention bottleneck, reaching 62.88% accuracy with improved training efficiency and performance comparable to BERT-based models
  • Built a custom preprocessing workflow for the 17K-tweet Sentimix dataset using Microsoft's Language Identification tool and transliteration-based lemmatization to resolve ambiguity in code-mixed Hindi-English text

Mentee

Microsoft Engage

Content-based recommendation system for a movie streaming platform

Summer 2022

Pythonscikit-learn
  • Built a content-based filtering system for movie recommendations using genre and cast features
  • Optimized the recommendation algorithm for a 25% increase in user engagement
  • Improved recommendation relevance and accuracy over the baseline approach

Live Demos — click and run, not just read about

Live demos

Other Projects — watch it work

Other projects

AGENTS

OIKOS — Household Financial AI

A conversational AI that acts as a household's financial nervous system — grounded in real Plaid transaction and budget data, it reasons with Gemini to answer questions by checking live budget context, then orchestrates flight/hotel search, restaurant discovery, and books the loop closed with Calendar events, a Gmail PDF itinerary, and Twilio confirmations to the whole family. Built for the Personal Executive Agent hackathon track. (Co-built with Krishna Sai)

Next.jsFastAPIGeminiPlaidSerpAPIGoogle APIsTwilioClerk
MULTIMODAL AI

ChalkTalk AI — Lecture Quality Auditor

A pedagogical copilot that analyzes lecture video recordings to tell active teaching (whiteboard, gestures, demonstrations) apart from passive teaching (scrolling through slides). Samples keyframes and high-motion segments with OpenCV, uses Tesseract OCR to distinguish static slide text from skewed handwriting, and fuses frame + audio context through Gemini 2.5 Flash to produce a time-stamped engagement heatmap and an AI-generated executive summary for the professor.

Next.jsFastAPIOpenCVTesseract OCRGemini 2.5 Flash
GENAI · AGENTS

Intelligent Multi-Agent Tutor

IntelliLearn Studio uses a hierarchical multi-agent architecture — a SessionCoordinator delegates tasks to specialized sub-agents that analyze study materials, generate quizzes, and adapt difficulty (Easy/Medium/Hard) in real time based on student performance. It identifies knowledge gaps and surfaces targeted resources (e.g. relevant YouTube links) to close them, and persists student progress and history across sessions using Google ADK's state management.

Google ADKGemini 2.5 FlashReactTailwind CSSshadcn/ui
GENAI

Socratic Checker

Built at DevFest DC 2026's build-a-thon. Given any topic, generates three diagnostic questions engineered to expose misconceptions rather than test recall, then evaluates free-text answers to return a diagnosis — the single concept the learner's understanding actually breaks on — plus one concrete next step. Where a quiz returns a score, this returns a root cause. (Co-built with Nikhil Arethiya)

Next.jsGemini API
REINFORCEMENT LEARNING

IntelliHVAC — DRL for HVAC Control

Forked and extended Google's Smart Buildings simulator to run a controlled SAC vs. PPO ablation — both agents trained in parallel on one GPU under identical compute budgets, optimizing a multi-objective reward balancing energy cost, carbon intensity, and ASHRAE comfort. Instrumented actor/critic loss straight from the TensorFlow graph to tie policy improvement to convergence rather than simulator variance. SAC outperformed PPO while completing 3x more training iterations.

PythonTensorFlowGoogle SBSimSACPPO
RESEARCH · RAG

Knowledge Graph vs. RAG for Code Generation

Constructed an 855K-node / 1.96M-edge AST-derived knowledge graph (NetworkX) from a large Python codebase and benchmarked it head-to-head against a FAISS vector index over 455K Python files for LLM-assisted code generation. Evaluated both retrieval strategies on LiveCodeBench, raising Pass@1 from a 64.5% baseline to 84.25% with the stronger approach, and diagnosed why graph traversal underperformed in isolation (77.50%) — keyword-dependent lookup struggles to match natural-language problem statements.

PythonNetworkXFAISSLiveCodeBench

Skills

// move your cursor through it

  • Python
  • PyTorch
  • LangChain
  • LangGraph
  • TensorFlow
  • FastAPI
  • Hugging Face
  • Multi-Agent
  • RAG
  • scikit-learn
  • Docker
  • Next.js
  • Git
  • C++
  • Vector DBs
  • Knowledge Graphs
  • LlamaIndex
  • Google ADK
  • SQL
  • LoRA
  • Prompt Engineering
  • AWS

Contact

// Interested in collaborating? Type 'help' to get started.

sathwik@portfolio: ~
Type 'help' to see available commands.
$

Or send a direct message