CP ClaimPilot HarnessAgent Reliability Lab

Open-source AI agent evaluation product

Can your claims Agent survive the messy evidence?

ClaimPilot turns policy exclusions, missing documents, conflicting evidence, privacy traps and prompt injection into repeatable tests. Pick a case, run an Agent, and inspect exactly why it passed or failed.

10adversarial cases
--risk patterns
--safe baseline score
--automated tests
Evaluation workspace
Interactive benchmark simulation · deterministic results

claim amount
Load evidence
Inspect risks
Score decision
Build replay
Ready. Select a case and agent, then run the evaluation.

Model Arena with evidence, not vanity scores.

每次实验都绑定数据集指纹、模型配置、运行时间、延迟和逐案 replay。规则基线与外部模型结果分开展示,不把模拟结果包装成真实模型成绩。

ExperimentLoading
Dataset fingerprintLoading
External model slotOpenAI-compatible · HTTP · command
ProfileRun typeAveragePass rateMedian latencyStatus
claimpilot arena cases --config benchmarks/models.json --require-allModel config ↗

Trace Arena: evaluate the path, not only the answer.

同一宗车险人伤案件,逐步比较 Agent 是否先查保单、读取证据、请求补件、升级人工,以及危险赔付动作是否经过审批。

AUTO-BODILY-INJURY-001

Cautious Trace

Human Review Workbench.

把 Agent 建议变成人工可确认、可修正、可导出的复核记录,为后续回归测试和反馈学习留下结构化证据。

Review queue
Agent recommendation
Human decision
Confirmed findings
Review statusNot reviewed0 saved reviews

From impressive Agent demo to measurable product reliability.

这是一个完整的 AI 产品闭环:从业务风险建模,到评测指标、交互体验、工程接入与持续回归,而不只是调用一次大模型 API。

01 · Product thesis

The problem is trust, not fluency.

理赔 Agent 真正的上线门槛不是“回答是否流畅”,而是在证据冲突、材料缺失和恶意输入下,能否做出安全、可解释、可复盘的下一步决策。

  • 定义核心用户:AI 产品、风控、理赔运营、算法工程
  • 把模糊的“可靠性”拆成可验证指标
  • 用 risky baseline 暴露系统性失败模式
02 · Evaluation design

Deterministic scoring

按业务结论、关键风险、补充材料、证据引用、禁止行为和注入防护打分,结果可重复、可进入 CI。

03 · Product experience

Replay before leaderboard

不仅告诉团队“得了多少分”,还展示 Agent 看到了什么、漏掉了什么,以及为什么失败。

04 · Integration

Adapter-first

支持 OpenAI-compatible、HTTP service 与 command adapter,真实 Agent 无需改动评测面。

05 · Delivery

Regression-ready

CLI、JSON artifacts、GitHub Actions、Pages 自动发布与自动化测试组成持续交付链路。

06 · Domain skill

Bodily injury claims processing

把人伤理赔中的事故因果关系、治疗时间线、医疗材料、误工损失与人工升级规则,转化为可复用 Skill 和可回归案例。

A small system with a production-shaped architecture.

同一份结构化 case pack 驱动本地 CLI、Agent 适配器、确定性评分、HTML replay 和 GitHub Pages 产品 Demo。

INPUTAdversarial case pack

Policy, claimant context, evidence, traps and expected safe behavior.

RUNTIMEAgent adapters

Built-in baselines, HTTP services, commands and OpenAI-compatible endpoints.

EVALScoring engine

Repeatable business and safety checks with CI quality gates.

OUTPUTReplay + artifacts

Human-readable decisions and machine-readable regression results.

See the whole product story in 60 seconds.

中文旁白演示从对抗案例出发,对比 100 分的证据优先轨迹与 5 分的危险捷径,并展示审批门禁、自动化测试和持续交付能力。

1080p · 60 seconds · Chinese narrationDownload MP4

Built as an AI product case study

ClaimPilot is the test range, not another claims chatbot.

我把这个项目作为一个真实产品问题来设计:如何帮助团队判断 Agent 是否具备上线条件,并让产品、业务、风控和工程看到同一份证据。

v0.3 已加入工具调用 Trace Arena、审批门禁、可追踪 Model Arena、人工复核闭环、10 个对抗案例与车险人伤 Domain Skill Pack;外部模型结果必须来自真实运行,缺失数据保持为空。