Customs and Compliance Agents
AI agents that prepare customs and compliance filings, with every number checked by code
- Role
- Led AI engineering in a small team at DigiSol
- Status
- Live
- Source
- Client work
- Stack
- Next.js 16TypeScriptshadcn/uiRadixTailwind 4TanStack TablePythonFastAPILangGraphAG-UIPostgreSQLLiteLLMLangfuseKeycloakoauth2-proxyCaddyDocker ComposeGitea CI
In numbers
135/ 135
amounts of a closed permit reproduced by the incentive-closure engine
1,185/ 1,185
cells of the client's own consumption matrix recomputed identically from raw parameters
166/ 166
capacity-report rows recomputed exactly and matched to the signed final reports
24/ 24
lines of an approved filing reproduced, kept as a permanent regression test
8,355
agent-side tests passed in one full run
85/ 85
import rows classified identically to the client's own hand-classified file
The problem
A large industrial manufacturer files three families of regulated documents each year: an annual production-capacity report, tax-incentive permit studies and customs permit applications. The rules lived in spreadsheets, screen-recorded meetings and the heads of a few specialists, and none of it was written down. A wrong figure in these documents has regulatory consequences, so a model that produces numbers on its own was unacceptable from the start.
The approach
I reverse-engineered the processes into written specs before any code, tagging each claim as proven, theory or open, with a proof script behind the proven ones. The order of work was fixed early: the generated documents had to be correct first, and the UI and the agent came after. I worked in a small team. The lead developer owned infrastructure, identity, the model gateway and releases; I owned the department modules, the validation of the calculation engines, the front-end contract and the governance patterns. I built my part with fleets of AI coding agents working under review gates I designed: I write the requirements and acceptance criteria, read the evidence and decide.
How it works
Engines held to the client's history
The model does no arithmetic. Each figure comes from a deterministic engine: renderers that reproduce the capacity report in each recipient's own format, an incentive-closure engine, a tariff-code classifier and a worksheet generator for permit applications. Before an engine was trusted, the original reference pages were vendored, their numeric output committed as golden JSON, and the port held digit-exact to it. The classifier is the one place a model judges, and only as the third tier, after a registered-link lookup and a rules engine built on the official tariff text, with an evidence pack in front of it.
Nothing is generated without review
Uploads are staged, diffed against the current state and matched row by row. An ambiguous match is refused and an unmatched name fails the ingest. The prior state is backed up before any write, a person reviews the data before a document can be generated, and generation runs from a frozen snapshot. A second, independent calculation then re-adds each generated file and discards it on any difference. Anything the engines cannot price shows an explicit no-pairing state instead of a number.
Page-first agents with signed approvals
Operators work on ordinary pages fed by REST data. Each department has one plain LangGraph StateGraph, a chat node and a tool node with one tool call per step, and no supervisor sits above them. Where an assistant helps, it sees the page state and drives the same controls a person would. Data-changing tools are risk-tiered: a medium-risk write parks on an approval card whose signature covers the canonical arguments, and a batch of edits becomes one approval, one audit decision and one revision bump. In an audit I found a path where front-end write tools changed the working database with no card at all, and deleted it.
Locked down inside the client's network
Production runs on a server inside the client's network behind single sign-on and a login gate. Each agent has its own identity, and its effective scope is the intersection of user, agent and resource permissions, so authority can only shrink. The application has no general internet access: outbound connections follow an allowlist with a data classification per destination, and a nightly egress audit checks that vendor telemetry stays off. Each AI chat is traced per user, each tool call, admin action and data write lands in an append-only audit table, and CI scans container images for vulnerabilities. The lead developer designed the identity seam; I wrote the outbound-connection requirements.
Release trains with review agents
Before a production release, seven review agents audited the candidate in parallel and filed their findings as issues. Fix agents worked on isolated worktrees, browser QA on the client's real files found new blockers in one flow, and the fixes were re-tested before the pull requests merged one CI run at a time behind a deploy runbook with a smoke test and a rollback step. Production then ran the fixed release. Migrations are forward-only, so the restore path is rehearsed.
What I chose, and what lost
Chose
Page-first: ordinary pages, with the agent embedded in them
Over
A chat-first portal where the agent is the main interface
Operators need pages they can read and check. Building the pages first and then letting the agent use them gives the agent the same controls a person has, and gives each change a surface someone can review.
Chose
One plain graph per department and no supervisor agent
Over
A supervisor that routes requests between department agents
A supervisor stands between the user and the tool call and blurs who decided what. One graph per department with one tool call per step keeps the audit trail readable.
Chose
A logic-free front-end kit: the browser never computes and never fetches
Over
Arithmetic and data fetching inside UI components
With one place that computes, there is one place to test against history. The boundary survived two full back-end replacements with a handful of changed lines.
Chose
A lean production profile with chat only on the classification page
Over
An assistant on each page
Where an agent added nothing over a deterministic page, I removed it. Fewer agents in production means fewer failure points and less to review.
Parts of the software, rebuilt to touch
The client's data never leaves its network, so these are recreations: the same anatomy, with invented products, people and numbers.
Exhibit 01
Approval card
The card an agent's data change parks on until a person accepts or rejects it.
Update consumption standards
Agent Kaloyan, acting for R. Petrova
3 edits, one approval
| Field | Current | Proposed |
|---|---|---|
| Line 7 base compound, kg per unit | 0.42 | 0.40 |
| Batch A-114 yield, % | 92.0 | 93.5 |
| Profile P-40 scrap allowance, % | 3.0 | 2.5 |
7f3a 91c2 e04b 5d18Exhibit 02
Dual-calculation sheet
The check that runs on each generated file: the committed figures beside an independent recomputation. A file with any difference is discarded.
Annual capacity report, section B
From snapshot S-0193, frozen at 09:40
| Line | Basis | Committed | Recomputed | Delta | Status |
|---|---|---|---|---|---|
| B.1 Line 7 output | 1,200 t × 0.92 | 1,104.0 | 1,104.0 | 0.0 | match |
| B.2 Line 9 output | 800 t × 0.95 | 760.0 | 760.0 | 0.0 | match |
| B.3 Batch A-114 | 400 t × 0.90 | 360.0 | 360.0 | 0.0 | match |
| B.4 Resin R4, imported | no price for the basis | No pairing | No pairing | no pairing | |
| B.5 Section total | B.1 to B.3 | 2,224.0 | 2,224.0 | 0.0 | match |
Exhibit 03
Replay match report
How an engine earns trust: its output replayed against a past filing that has an answer key, with the kind of evidence stated.
119/ 120 matched
Tolerance: exact, to the cent
Replayed against a past filing with an answer key.
Mismatches
| Row | Expected | Produced | Traced cause |
|---|---|---|---|
| Row 47 | 1,250.00 | 1,200.00 | The source spreadsheet adds row 47 twice; the filed figure carries the error. |
golden: permit_closure_case_aExhibit 04
Staged import diff
What a person sees between uploading a workbook and changing any data.
rates-update.xlsx
Sheet Rates, 64 rows read, nothing written yet
New2
- Profile P-52: 4.10 per m
- Batch A-120: 91.0 %
Changed4
- Line 7 base compound: 0.42 to 0.40
- Profile P-40: 3.85 per m to 3.90 per m
- Resin R4: 2.20 per kg to 2.35 per kg
- Batch A-114: 92.0 % to 93.5 %
Unmatched1
- Profil P-44: no target with this name
An unmatched name fails the whole ingest. Fix the file and upload it again.
Exhibit 05
Audit trail
The append-only record of AI tool calls, resumes, memory writes, workbench changes and admin actions. It stores references and no document content.
Audit trail
References only, no document content
| Time | Actor | Action | Target | Outcome | Trace |
|---|---|---|---|---|---|
| 09:40:12 | Agent Kaloyan for R. Petrova | Tool call | propose_edits, 3 edits | parked on card C-77 | |
| 09:41:03 | R. Petrova | Approval | card C-77 | accepted | person |
| 09:41:03 | Agent Kaloyan for R. Petrova | Data write | standards, revision 41 to 42 | written | |
| 09:44:50 | Agent Desislava for N. Stoyanov | Resume | session S-12 | resumed | |
| 09:46:18 | Agent Desislava for N. Stoyanov | Memory write | note M-9 | stored | |
| 09:52:31 | D. Ilieva | Admin | role grant, page Imports | granted | person |
| 10:02:07 | Agent Boyan for R. Petrova | Tool call | generate_report from S-0193 | accepted after re-check |
Outcome
The system is live in production, built on four department agents and one pipeline for three document families. Correctness is shown by replay: the incentive-closure engine reproduces all 135 amounts of a closed permit, the client's own consumption matrix recomputes identically in all 1,185 cells, 166 of 166 capacity rows match the signed final reports, the application generator reproduces all 24 lines of an approved filing, and import-row classification matched the client's hand-classified file in 85 of 85 rows. In its latest full runs the agent-side suite passed 8,355 tests and the web suite 1,682; line coverage stood at 92.1 percent against an enforced floor. The same checking caught an overstated accuracy claim before it went out: an independent re-derivation did not support it, and it came out of the client draft. Limits: one of the four department agents is disabled in production, and error tracking was not yet reporting from production at the last check, so I claim tracing, token accounting and audit only.
What comes next
Per-page entitlements, so a user holds only the pages their job needs, and error tracking switched on in production.