Cedars Digital · May 2024 – Present

AI Code Review in the MR Pipeline

Every MR on our GitLab main pipelines gets reviewed against its own service's standards, and the model behind it is swappable

0
Code changes to onboard a service
2
Internal survey rounds
11+
Years of architecture reuse

Problem

Review throughput stopped keeping up with the code the engineers were producing. What bothered me wasn’t the queue length, it was where senior attention went: naming, missing tests, config that had drifted since the last release. The architecture questions, the ones only they could answer, waited behind all of that.

Architecture

flowchart LR
  MR[GitLab MR trigger] --> Loader["Profile Loader<br/>Service Registry"]
  Loader --> Standards["shared + system-specific<br/>review standards"]
  Standards --> Retrieval["Retrieval layer<br/>hybrid search + token budget<br/>(fail-open)"]
  Retrieval --> Composer[Prompt Composer]
  Composer --> AI["LLM Provider<br/>(pluggable)"]
  AI --> Reflect["Self-reflection validation<br/>(optional)"]
  Reflect --> Comment[Post MR comment]

A service joins by registering a profile. Nothing lands in the service repo, so onboarding costs a team zero code changes and the review standards stay in one place instead of being copy-pasted into every pipeline.

Retrieval, not a bigger prompt (2026)

The first version handed the reviewer everything it might need. That works until the review standards, the API contracts and a year of accumulated review history all want room in the same context window, and then it just stops. So in 2026 I replaced context assembly with retrieval, and built the vector store rather than adopting one: SQLite for storage, 768-dimension embeddings, brute-force L2 in plain JS. At this corpus size an external vector database is operational weight I’d be carrying for nothing.

Retrieval is hybrid. BM25 and embedding search get fused with RRF and routed in two structured tiers, and each source has its own token budget so one corpus can’t crowd the others out of the prompt. A weekly offline job builds the index and publishes it atomically with a sha256 checksum that the consumer verifies before switching over. Every stage fails open: hybrid degrades to lexical, lexical degrades to loading everything. A broken index makes a review slower. It never blocks a merge.

Where the benchmark actually lands: on recall, hybrid beats lexical-only on all three corpora. Ranking is a weaker story. By MRR only the contract corpus improved clearly, and the other two trade places with lexical search depending on the query. What I’ll defend is that the right context gets found. Not yet that it gets ranked better. This layer has been live for weeks, not quarters, so read it as an early result.

My role

Design of v1.0 through the v2.0 expansion, including the profile module, the self-reflection pass, and the rollout across microservices.

Impact

Lessons — the same skeleton, three times over

In 2015 at iPanSec I built A4P: a Python subprocess driving MobSF, a crawler scraping the report, structured output at the end.

In 2024 at Cedars, AI Code Review swapped MobSF for an LLM API. The rest of the skeleton barely moved.

In 2026 the retrieval layer went in and the skeleton still barely moved. The one real change is that the model now gets a selection instead of the whole pile.

A senior engineer’s long-term value isn’t the newest framework they can name. It’s noticing which old problem just got a better solution.