AI Pre-Audit + Human Adjudication · Clear 300 Pending Reimbursements in 5 Days
Built for finance departments in precision-manufacturing and other large enterprises, the system uses the paradigm of "agent pre-audit + risk tiering + human adjudication": AI takes over the four repeated steps — invoice verification, duplicate detection, standard matching, and image check — while the finance manager keeps the final adjudication right. We ran a 6-day PoC for a 480-person manufacturing client and cleared all 300 pending reimbursements within 5 business days: manual review workload dropped 76%, average handling time fell from 15 to 4 minutes, and every abnormal judgment came with evidence — 100% explainable.




AI does not replace finance accountability — it takes over the standardized, repeatable labor. Audit shifts from "investigate four steps from scratch" to "adjudicate on Agent-supplied evidence."
Stay on the single line of "reimbursement audit" — no ERP rebuild, no post-audit flow, no non-reimbursement work, no auto-payment. Not creeping the scope is the prerequisite for 6-day delivery.
Low-risk auto-pass (manager spot-check), medium/high routed to the human review workbench. Reviewers shift from "look up four steps" to "adjudicate on Agent-supplied evidence".
Every abnormal judgment must attach evidence fragments and rule basis. Managers can send the rejection reason to employees directly without further explanation — explainability is the prerequisite for finance to trust the system.
The whole stack runs on the client's virtualized in-house resources: data never leaves the network, rules are configurable, ERP is read-only via API — no intrusion into the core ledger, simple ops.
FDE cross-validated the pain points through system data, policy materials, and 7 stakeholder interviews during the PoC. The pain points were not directly given by the client.
Reviewers hand-run "verify → check amount → check standard → check duplicate", ~15 min per claim. 300 claims ≈ 9–10 person-days, chronic backlog.
No verification API, reviewers judge by experience. Sampling found 3 cases of abnormal invoice codes — risk of fake/cloned invoices slipping through.
Whether the same invoice/amount has been reimbursed depends on memory and Excel cross-check. Cross-validation found 11 suspected duplicates.
Travel/entertainment standards conflict in interviews: B says "lodging ≤400/night", D says "tier-1 ≤500, others ≤350". Policy files are missing; same case, different decisions.
Of 4,222 images, 217 lack claim linkage or are unclear. Manual flipping is easy to miss, with no system prompt.
300 claims mix high-risk (large amount, over-standard) and low-risk (small amount, complete docs) in one queue — resources misallocated.
Rejection reasons are mostly verbal/abbreviated, lacking structured evidence. Employees dispute, managers cannot audit.
Designed around the full chain "ingest → pre-audit → tier → adjudicate → feedback". Every capability has clear priority and boundaries.
Pull claims, invoice ledger, and receipt images via open-platform API; build claim–invoice–image linkage (F1).
Recognize seller, amount, tax ID, invoice number, and date from images; output structured fields (F2).
Cross-check invoice number, code, and amount against ledger and OCR; if tax-authority API available do real-time verify, otherwise rule-approximate with confidence label (F3).
Flag second occurrence of same invoice number / same amount + same seller. Cross-checked 11 suspected duplicates in the ledger (F4).
Validate lodging, meals, transport upper limits by type. Rules visual + configurable, finance team self-service (F5).
Verify claim department matches budget attribution and does not exceed department budget (F6).
Output low/medium/high risk; low risk auto-passes, medium/high goes to human review (F7).
List claims to review with Agent-flagged items, evidence fragments, and suggested conclusions; single-screen review (F8).
Rejections must include structured reasons; all judgments leave a trace and are exportable (F9).
Real-time display of throughput, auto-pass rate, abnormal recall, manual effort, and more (F10).
Employee submits → system ingest layer (API + OCR) → agent pre-audit engine → risk tiering → low-risk auto-post / medium-high human adjudication → conclusion feedback + metrics dashboard.
Based on 6-day PoC measurements: baseline = pre-PoC fully-manual; target = post-PoC AI pre-audit + human adjudication.
The difficulty is not single-point rules but connecting discover–tier–adjudicate–feedback into an auditable, quantifiable, reusable pipeline.
Stuck to the single line of "reimbursement audit" for 6 days. No ERP rebuild, no post-audit flow, no non-reimbursement work, no auto-payment. Not creeping the scope is what made 6-day delivery possible.
Not "AI replaces human review" but "AI frees humans from repeatable labor": low-risk auto-passes (manager spot-check), medium/high enters the workbench, reviewers shift from "look up four steps" to "adjudicate on Agent-supplied evidence".
Every abnormal judgment must attach evidence and rule basis. Managers can send the rejection reason to employees directly without further explanation — explainability is the prerequisite for finance to trust the system.
The whole stack runs on the client's in-house virtualized resources. Data never leaves the network, ERP is read-only via API, rules are visual + configurable. 300-claim batch pre-audit ≤ 4 hours.
Convert technical capability into quantifiable, reusable business value.
Product content has been published based on internal materials. The following areas are planned for further development:
Explore Xianma AI solutions in other domains