Skip to content
Selected work

Case study 03

AI Code-Review Orchestrator

A webhook-driven agent that reviews every pull request across multiple repositories before a human does - and cut per-PR review time from 15-20 minutes to about 5.

Outcome

15-20 min → ~5 min per pull request

Before
15-20 min
After
~5 min
Coverage
Every repository

Context

Built in-house at Cygnis Media, where I work. Human review had become the team's bottleneck, so I designed and shipped the orchestration layer that turns incoming pull-request webhooks, across every repository, into reviewed, merge-gated code before an engineer ever opens the PR.

The problem

Human review was the bottleneck: 15-20 minutes per pull request, multiplied across every repository and every engineer's day. The goal was never to replace reviewers - it was to make sure that by the time a human opened a PR, the mechanical findings were already on the table, with a defensible signal on whether the merge should be gated at all.

The outcome

For the business: review time per pull request dropped from 15-20 minutes to about 5, across every repository the orchestrator watches. For the system: an event-driven review pipeline with whole-repository context, fair scheduling, stale-work cancellation, explainable merge gating, and an evaluation harness that keeps the agent honest.

The approach

  1. Whole-repo context over diff-only review. The agent reads each change against the entire repository, not just the patch - so it catches architectural drift, inconsistent patterns, and violations of conventions the rest of the codebase already follows. A diff read in isolation can see none of that.

  2. Fairness before throughput: a per-repository FIFO queue with round-robin dispatch. I chose per-repo fairness over a single global queue because one busy repository must never starve review for the others.

  3. SHA coalescing and stale-run abortion: a new push supersedes any in-flight review of the outdated commit, so compute is only ever spent on code that can actually merge.

  4. Merge gating on three independent signals - severity, risk, and confidence. A finding blocks a merge only when all three justify it. I chose three explicit signals over a single blended score because a gate you can't explain is a gate teams learn to override.

  5. Quality is measured, not asserted: 13 vitest suites plus a golden-set evaluation harness that scores review output against known-good judgments.

Built with

  • AI agents
  • Multi-repo orchestration
  • Express
  • Redis
  • Evaluation harness