IISc · MULTI-AGENT · 2026
Multi-Agent Code Review & Auto-Debugging System
A LangGraph multi-agent system that reviews a GitHub pull request with four agents in parallel (security, bugs, style and performance) and commits fixes back to the pull request. Built in the IISc Executive Education Program in Agentic & Generative AI.
- Context
- IISc Bengaluru · Executive Education Program – Agentic & Generative AI
- My role
- Architected the multi-agent platform; built the RAG-powered code intelligence system
- Year
- 2026
- Stack
- LangGraph · FastAPI · ChromaDB · Claude · React
Context
I built this system in the Executive Education Program in Agentic & Generative AI at the Indian Institute of Science (IISc) Bengaluru, which I completed in 2026. The goal was to automate secure code review and AI-assisted debugging workflows.
My role
I architected the multi-agent platform and built the RAG-powered code intelligence system behind it.
How it works
A review starts from a pull-request URL or a GitHub webhook. The FastAPI service accepts it at once (202 Accepted), runs the review in the background and streams progress to a React dashboard.
- Supervisor–Worker orchestration. A LangGraph orchestrator acts as the supervisor. It fetches the pull request's changed source files through the GitHub REST API, then fans the work out to four analysis agents that run in parallel: security, bug detection, style and performance.
- Static checks first, then the LLM. Each agent pairs a cheap, deterministic check with an LLM (Claude, at temperature 0). The security agent retrieves the most relevant OWASP and CWE references from a ChromaDB vector store (RAG). The bug and performance agents run Python AST analyzers, and the style agent runs Ruff. Their results go to the LLM as hints, so the model spends its effort on what static checks can't see.
- One findings contract. Every agent returns findings in the same JSON shape: severity, confidence, line range, suggestion and CWE ID. LangGraph reducers merge the parallel results into one review.
- Patch generation. If any finding is medium severity or higher, a fix agent asks the LLM for corrected code, checks that the code still compiles, and commits the fixes to the pull request's own branch: one commit per category.
- Test stage. The graph ends with a test stage that hands back the pull-request link. In the current build, generating tests there is switched off.
Key decisions and trade-offs
1. Parallel agents, merged with reducers
Decision. The four analysis agents run at the same time and write to shared state through LangGraph reducers.
Trade-off. A review takes about as long as its slowest agent, not the sum of all four: typically 30–60 seconds, depending on the size of the pull request and the model's latency. The cost is that concurrent writes need explicit merge rules.
2. Static analysis as hints for the LLM
Decision. The AST analyzers and Ruff run before the LLM, and their findings go into its prompt. Duplicate style findings are removed.
Trade-off. Deterministic checks are fast and never hallucinate, and they let the model focus on deeper issues. However, the AST analyzers cover Python only, so for other languages the agents rely on the LLM alone.
3. RAG that degrades gracefully
Decision. The security agent grounds its review in OWASP and CWE references retrieved from ChromaDB. If the vector store is unavailable, it falls back to a built-in set of core references.
Trade-off. The fallback is smaller, so findings can be less specific, but a review never fails because the vector store is down.
4. Commit through the GitHub API, one commit per category
Decision. There is no local clone. Each fix is committed with the Git Data API (blob, tree, commit, then a ref update) in a fixed order: security, bugs, style, performance.
Trade-off. Each category can be reviewed or reverted on its own, and there is no clone time or disk to manage. In return, a compile check is the only safety net before a commit. The scope is capped at medium-severity findings and above, and at 10 files per category. That is why the fixes land on the pull request for a person to review, never straight on the main branch.
5. In-process first, a queue later
Decision. The first version runs each review as a background task inside the API process. A Redis/ARQ worker is already designed for the next step.
Trade-off. The system is simple to run and debug, but a restart loses any review in flight until reviews move to the queue.
Reliability and security
- GitHub webhooks are verified with HMAC-SHA256 signatures, and the API is rate-limited per client.
- One failing file never stops an agent. Typed errors carry their own retry policy: for example, exponential back-off for GitHub and LLM errors, and no retry when a prompt is too long.
- Correlation IDs run through structured logs, and LangFuse traces the agent and LLM calls.
Evaluation
I evaluated agent performance across four scenarios: code review, bug fixing, patch generation and test generation. The evaluation combined automated evaluation metrics, test execution and RAGAS-based evaluation.
Stack
Python · LangGraph · LangChain · Claude · ChromaDB (RAG over OWASP / CWE) · FastAPI · server-sent events · React · Redis / ARQ · LangFuse · Docker · RAGAS