Case study 03
AI Code-Review Orchestrator
A webhook-driven agent that reviews every pull request across multiple repositories before a human does - and cut per-PR review time from 15-20 minutes to about 5.

Outcome
15-20 min → ~5 min per pull request
- Before
- 15-20 min
- After
- ~5 min
- Coverage
- Every repository
Context
Built in-house at Cygnis Media, where I work. Human review had become the team's bottleneck, so I designed and shipped the orchestration layer that turns incoming pull-request webhooks, across every repository, into reviewed, merge-gated code before an engineer ever opens the PR.
The problem
Human review was the bottleneck: 15-20 minutes per pull request, multiplied across every repository and every engineer's day. The goal was never to replace reviewers - it was to make sure that by the time a human opened a PR, the mechanical findings were already on the table, with a defensible signal on whether the merge should be gated at all.
The outcome
For the business: review time per pull request dropped from 15-20 minutes to about 5, across every repository the orchestrator watches. For the system: an event-driven review pipeline with whole-repository context, fair scheduling, stale-work cancellation, explainable merge gating, and an evaluation harness that keeps the agent honest.
The approach
Whole-repo context over diff-only review. The agent reads each change against the entire repository, not just the patch - so it catches architectural drift, inconsistent patterns, and violations of conventions the rest of the codebase already follows. A diff read in isolation can see none of that.
Fairness before throughput: a per-repository FIFO queue with round-robin dispatch. I chose per-repo fairness over a single global queue because one busy repository must never starve review for the others.
SHA coalescing and stale-run abortion: a new push supersedes any in-flight review of the outdated commit, so compute is only ever spent on code that can actually merge.
Merge gating on three independent signals - severity, risk, and confidence. A finding blocks a merge only when all three justify it. I chose three explicit signals over a single blended score because a gate you can't explain is a gate teams learn to override.
Quality is measured, not asserted: 13 vitest suites plus a golden-set evaluation harness that scores review output against known-good judgments.
Built with