AI/ML Engineer · Chicago, IL

Jaya Prakash Yadav Gorla

I build LLM, RAG and multi-agent systems, from data pipelines and fine-tuning to evaluation and deployment.

Now
AI Data Scientist at Taisho Systems, building AI features for MCP Tutor, an AI learning platform
Education
M.S. Artificial Intelligence, DePaul University (2026)
Jaya Prakash Yadav Gorla in a navy suit, walking across a footbridge

Projects

Open “How it works” on a project for its architecture and how each result was measured.

  1. MediQuery: Multi-agent medical QA

    2026 · Co-authored with Kunal Tamhane

    Medical question answering that stops instead of guessing when no source comes back.

    Manuscript in preparation

    How it works for MediQuery
    Problem
    Language models answer medical questions confidently even when no evidence supports the answer.
    Approach
    A LangGraph supervisor orchestrates three retrieval agents (PubMedBERT embeddings in Qdrant, NCBI, web search), and answers are generated through Amazon Bedrock. We also fine-tuned Llama 3 8B with 4-bit QLoRA on PubMedQA. A circuit breaker returns an insufficient-evidence reply instead of guessing when no agent finds a source.
    Evaluation
    Pilot: the full pipeline scored 0.92 BERTScore-F1 on 5 PubMedQA questions.
    MediQuery architecture A medical question goes to a LangGraph supervisor, which queries three retrieval agents: PubMedBERT with Qdrant, NCBI, and web search. A circuit breaker then either returns an insufficient-evidence reply when no source was found, or lets Amazon Bedrock generate an answer with sources. Medical question LangGraph supervisor PubMedBERT + Qdrant NCBI Web search Circuit breaker Insufficient-evidence reply (no source) Answer with sources generated via Bedrock
    Simplified architecture.
    Result

    3 retrieval agents + circuit breaker

    Stack
    • LangGraph
    • Amazon Bedrock
    • Qdrant
    • PubMedBERT
    • Llama 3 8B
    • QLoRA
  2. Multimodal AML detection: Graph ML for anti-money laundering

    2025–26 · DePaul MLOps course, team of 4

    Flags money laundering by fusing a transaction graph, behavioral sequences and synthetic payment-memo text.

    On a 4-person team, I owned the GraphSAGE encoder, late-fusion head and SHAP explanations.

    How it works for Multimodal AML detection
    Problem
    Laundering signals are spread across transaction networks, payment text and account behavior, so any single signal misses some of them.
    Approach
    Three encoders (GraphSAGE, BiLSTM, DistilBERT) feed a late-fusion MLP with Platt calibration and SHAP explanations. The team shipped it with MLflow, DVC, Docker and a Cloud Run CI/CD pipeline.
    Evaluation
    Elliptic Bitcoin graph (203,769 transactions, 46,564 labeled). Test AUC-PR: GraphSAGE 0.93; GraphSAGE + BiLSTM 0.947; adding DistilBERT on synthetic memos 0.9975, against 0.989 for an XGBoost baseline on tabular features. The memos are generated from each transaction's label, so the fused score shows the pipeline works end to end, not real-world accuracy. Halving the GraphSAGE hidden size (256 to 128) made epochs 1.92× faster (0.69 s to 0.36 s); the 0.93 model uses 256.
    Multimodal AML architecture Three encoders feed a late-fusion MLP: GraphSAGE on the transaction graph, a BiLSTM on behavior sequences, and DistilBERT on synthetic payment memos. The fused output is calibrated with Platt scaling to give a risk score, and SHAP explanations accompany it. The GraphSAGE encoder, the fusion MLP and the SHAP explanations are the parts I built. Transaction graph GraphSAGE Behavior sequences BiLSTM Synthetic memos DistilBERT Late-fusion MLP Platt calibration Risk score SHAP explanations Blue outline: the parts I built
    Simplified architecture.
    Result

    0.93 AUC-PR my GraphSAGE branch, Elliptic Bitcoin graph

    Stack
    • PyTorch Geometric
    • DistilBERT
    • BiLSTM
    • SHAP
    • MLflow
    • DVC
    • Docker
    • Cloud Run
  3. Tiny Dreamer: World-model RL for driving

    2026 · DePaul course project

    A Dreamer-style agent that learns to drive inside its own imagined rollouts on highway-env.

    How it works for Tiny Dreamer
    Problem
    Learning to drive from pixels with model-free RL takes a large number of real environment steps.
    Approach
    Built from scratch in PyTorch: a CNN encoder, an RSSM world model and an actor-critic trained on imagined rollouts, with CAPS smoothness regularization. Trained on about 184K agent steps.
    Evaluation
    Best checkpoint (cycle 20, 10 episodes): 86.4% lane-keeping, measured as the share of time on the road. Final policy over 20 episodes: 79.6% lane-keeping, off-road rate 20.4% against 32.7% for a random policy.
    Tiny Dreamer architecture A 64 by 64 camera frame passes through a CNN encoder into an RSSM world model with a 512-unit GRU state and a 32-dimensional stochastic state. Twenty-step imagined rollouts in that latent space train an actor-critic, which outputs steering and acceleration. 64×64 frame CNN encoder RSSM world model GRU 512 + 32 stochastic Imagined rollouts 20 steps in latent space Actor-critic Steer and accelerate
    Simplified architecture.
    Result

    86.4% lane-keeping share of time on the road, best checkpoint, 10 episodes

    Stack
    • PyTorch
    • highway-env
    • Gymnasium
    • RSSM
    • Actor-critic
    • CAPS
  4. Agentic RAG: Cleantech research Q&A

    2025 · DePaul NLP course project

    A tool-using research agent over 20,000+ cleantech articles, with a guardrail that reviews its answers.

    How it works for Agentic RAG
    Problem
    Cleantech research questions often need several documents plus current scholarly sources to answer well.
    Approach
    A LangChain tool-calling agent (GPT-4o mini) over a ChromaDB index of 20,111 documents (133,458 chunks). I extended a baseline agent (retriever and summarizer) with an OpenAlex scholarly-search tool and an answer-review guardrail.
    Evaluation
    A GPT-4o mini judge (the same model the agent runs on) scored 50 benchmark answers against reference answers: 44 rated 4 of 5, 4 rated 5 of 5, and 2 rated 3. On 23 questions, ROUGE-L F1 rose from 0.104 to 0.126 once the baseline gained OpenAlex search and the guardrail.
    Agentic RAG architecture A question goes to a tool-using GPT-4o mini agent, which calls three tools and receives their results: a ChromaDB retriever over 133,458 chunks, a summarizer, and OpenAlex scholarly search. An answer-review guardrail then accepts or revises the agent's answer before it is returned. Question Tool-using agent GPT-4o mini Answer review accept or revise Answer Retriever ChromaDB, 133,458 chunks Summarizer OpenAlex search
    Simplified architecture.
    Result

    96% of answers rated 4 or 5 of 5 GPT-4o mini judge, 50 questions

    +21% ROUGE-L F1 vs the baseline agent, 23 questions

    Stack
    • LangChain
    • ChromaDB
    • GPT-4o mini
    • sentence-transformers
    • OpenAlex API

Also built

  • Robotics coursework: control, planning and estimation 2026 · DePaul course

    PID control in Drake and MeshCat, direct-transcription trajectory optimization, Kalman filtering, RRT and GCS-inspired planners, and a SAC agent on MuJoCo HalfCheetah.

  • Airbnb price prediction 2025 · group project

    48,895 New York City listings: temporal and geospatial features, LightGBM and XGBoost models, SHAP explanations and a borough-level fairness check.

Experience

  1. – Present

    Taisho Systems, AI Data Scientist

    Wisconsin, USA · mcptutor.com

    • Build AI and data science features for MCP Tutor, an AI learning platform delivered to AI assistants through the Model Context Protocol.
    • Support client projects with AI and data science solutions.
  2. –

    DePaul University, Data Scientist

    Chicago, IL

    • Cut manual processing time by 85% by automating ingestion of 100,000+ JSONL records across 3 research projects.
    • Designed innovation metrics and NetworkX centrality models to evaluate 50,000+ biopharma patents, identifying 5 innovation hubs and 3 emerging clusters; findings supported 2 funded grant proposals.
    • Raised insight-extraction efficiency by 40% and cut analysis errors by 30% by standardizing PCA and temporal-analytics pipelines across 3 multi-year studies.
    • Python
    • SQL
    • Semantic ETL
    • NetworkX
    • PCA
  3. –

    AriesView, AI Research Intern

    Boston, MA

    • With legal experts and 2 engineers, built and deployed an OCR-based RAG system processing 1,000+ legal documents a week, a 40% accuracy gain.
    • Added vector search with metadata filtering, reaching 85% accuracy (F1 0.83) and cutting manual review time by 60% for 5 reviewers.
    • Tuned PostgreSQL vector indexes and CI/CD pipelines, cutting query latency by 25% while scaling to 10,000+ daily queries.
    • LangChain
    • FAISS
    • PaddleOCR
    • PostgreSQL
    • CI/CD
  4. –

    Cognizant, Junior Data Scientist

    India

    • Lifted predictive model accuracy by 35% across 2 business initiatives with feature pipelines built on 200+ complex SQL queries.
    • Promoted from trainee after exceeding performance baselines by 20%.
    • Led end-to-end model validation and release coordination across data and engineering teams, with zero failed production deployments.
    • SQL
    • Feature engineering
    • Model validation

About

I'm an AI/ML engineer with 3 years of industry and research experience, across data science, RAG and multi-agent systems.

I like working where research meets engineering: taking an approach that works in a notebook and turning it into a system with clear evaluation behind it.

Skills

Languages
Python, SQL, C++, Bash
GenAI & LLMs
RAG, multi-agent systems, LangChain, LangGraph, LlamaIndex, Hugging Face Transformers, LoRA/QLoRA fine-tuning, prompt engineering, LLM evaluation, FAISS, ChromaDB, Qdrant
Machine learning
PyTorch, TensorFlow, scikit-learn, XGBoost, LightGBM, GNNs (GraphSAGE), BERT-family models, OpenCV, reinforcement learning
MLOps & cloud
Docker, Kubernetes, MLflow, DVC, GitHub Actions CI/CD, FastAPI, PostgreSQL, AWS (Bedrock, SageMaker), GCP (Vertex AI, Cloud Run), Azure AI

Education

  • DePaul University

    M.S., Artificial Intelligence · Chicago, IL · 2024–2026

  • Siddharth Institute of Engineering & Technology

    B.Tech, Electrical, Electronics and Communications Engineering · Puttur, India · 2018–2022

Certifications & courses

  • Google Cloud Gen AI Professional

  • Hugging Face AI Agents Fundamentals

  • AWS Educate Machine Learning Foundations

  • AWS Educate Introduction to Generative AI

Contact

Open to conversations about LLM systems, RAG and applied ML. Email is the fastest way to reach me.

Send a message

All fields are required.