ACE-Style Agent Evaluation and Playbook Promotion Report
Contract Review Agent — Evaluation Report · effora-ace-demo
Passed
ACE-style interview / internal experiment prototype. Not H.I.G. proprietary ACE. Not an MCP server. Controlled mock scores are not a production readiness benchmark.
Score
75/100Before
────▶
100/100After
+25 points · scale 0-100
Decision
Promote playbook v46Review
Human review not requiredKey system learning
When reviewing contracts, extract termination, indemnity, liability, and payment risks; cite document snippets for each risk claim; explicitly list missing evidence; and call out conflicting liability or termination language across documents.
Validation summary
- ✓ Required risk categories covered (or gaps stated)
- ✓ Citations present in output
- ✓ Groundedness / missing-evidence language check
- ✓ Missing-evidence instruction present after learning
- ✓ Cross-document conflict guidance present after learning
- ⚠ Sample size and test coverage limited (n=3 sample documents; single controlled demo task)
Promoted artifact
Version: v46
Path: /app/data/playbooks/playbook_v46.json
Reason: score improved
Run trace
Baseline: run_56b241cacf
Candidate: run_e257a016c9
Evaluated: July 29, 2026
Evaluation dimensions (candidate run)
| Task Completion | True |
|---|---|
| Citation Presence | True |
| Completeness | True |
| Groundedness | True |
| Output Format Compliance | True |
| Latency Ok | True |
| No Failures | True |
| Hallucination Rate | Not separately measured in this prototype |
| Security | Not assessed beyond demo logging hygiene |
| Authorization Behavior | Not assessed (no auth on public demo endpoint) |
| Token Usage | Not metered (ACE_MOCK deterministic path) |
Field semantics
Technical appendix — JSON
{
"evaluation_metadata": {
"service": "effora-ace-demo",
"mock_mode_env_ACE_MOCK": "1",
"scoring_scale": "0-100"
},
"added_instructions": [
{
"id": "pb_d91bbce0be",
"instruction": "When reviewing contracts, extract termination, indemnity, liability, and payment risks; cite document snippets for each risk claim; explicitly list missing evidence; and call out conflicting liability or termination language across documents.",
"source_run_id": "run_56b241cacf",
"confidence": 0.78,
"created_at": "2026-07-29T23:12:30.077401+00:00",
"updated_at": "2026-07-29T23:12:30.077406+00:00",
"helpful_count": 0,
"harmful_count": 0,
"status": "active",
"human_review_required": false
}
],
"removed_instructions": [],
"changed_instructions": [],
"run_ids": {
"baseline": "run_56b241cacf",
"candidate": "run_e257a016c9"
},
"promotion": {
"promoted": true,
"version": "v46",
"path": "/app/data/playbooks/playbook_v46.json",
"reason": "score improved",
"before_score": 75,
"after_score": 100
},
"raw_response": {
"service": "effora-ace-demo",
"mode": "1",
"mock_mode": true,
"task": "Review the sample contracts and identify key contract risks with citations, including termination, indemnity, liability, and payment terms. Note missing or conflicting evidence.",
"before_score": 75,
"after_score": 100,
"score_scale": "0-100",
"promotion": {
"promoted": true,
"version": "v46",
"path": "/app/data/playbooks/playbook_v46.json",
"reason": "score improved",
"before_score": 75,
"after_score": 100
},
"what_system_learned": "When reviewing contracts, extract termination, indemnity, liability, and payment risks; cite document snippets for each risk claim; explicitly list missing evidence; and call out conflicting liability or termination language across documents.",
"diff": {
"added": [
{
"id": "pb_d91bbce0be",
"instruction": "When reviewing contracts, extract termination, indemnity, liability, and payment risks; cite document snippets for each risk claim; explicitly list missing evidence; and call out conflicting liability or termination language across documents.",
"source_run_id": "run_56b241cacf",
"confidence": 0.78,
"created_at": "2026-07-29T23:12:30.077401+00:00",
"updated_at": "2026-07-29T23:12:30.077406+00:00",
"helpful_count": 0,
"harmful_count": 0,
"status": "active",
"human_review_required": false
}
],
"removed": [],
"changed": []
},
"run1_id": "run_56b241cacf",
"run2_id": "run_e257a016c9"
}
}
Machine-readable: /demo.json · Health: /health