Exam CCAR-F Topic 1 Question 118 Discussion
Actual exam question for Anthropic's CCAR-F exam
Question #: 118
Topic #: 1
Question #: 118
Topic #: 1
You are building a structured data extraction system using Claude. The system extracts information from unstructured documents, validates the output using JSON schemas, and maintains high accuracy. It must handle edge cases gracefully and integrate with downstream systems.
After deployment, you find that 12% of extractions contain semantic errors that pass JSON Schema validation--for example, a duration such as "30 minutes" is incorrectly placed in an ingredient-quantity field. Human reviewers have the capacity to check only 20% of extractions.
Which approach most effectively allocates reviewer attention?
After deployment, you find that 12% of extractions contain semantic errors that pass JSON Schema validation--for example, a duration such as "30 minutes" is incorrectly placed in an ingredient-quantity field. Human reviewers have the capacity to check only 20% of extractions.
Which approach most effectively allocates reviewer attention?
Suggested Answer: A Vote an answer
Option A uses the limited review capacity as a risk-ranking mechanism rather than spending it uniformly. JSON Schema validation can prove that an output has the expected types and structure, but it cannot detect a semantically misplaced value such as a duration stored as an ingredient quantity. Field-level confidence can identify the exact parts of a record that warrant inspection, provided those scores are calibrated against labeled examples rather than trusted at face value. The labeled validation set must mirror the production distribution and include edge cases; Anthropic's evaluation guidance explicitly recommends task-specific evaluations, representative distributions, measurable criteria, and automated grading where possible.
Calibration then converts raw confidence into an empirically tested estimate of error risk, allowing reviewers to inspect the lowest- confidence records or fields first.
Calibration then converts raw confidence into an empirically tested estimate of error risk, allowing reviewers to inspect the lowest- confidence records or fields first.
by Drew at Oct 08, 2026, 05:58 AM
0
0
0
10
Comments
Upvoting a comment with a selected answer will also increase the vote count towards that answer by one. So if you see a comment that you already agree with, you can upvote it instead of posting a new comment.
Report Comment
Commenting
You can sign-up / login (it's free).