Document field extraction
Extracts a defined schema from mixed business documents into a normalized, source-cited table while preserving original wording, missing fields, and confidence notes. Use for invoices, forms, receipts, applications, statements, certificates, and document registers.
Veröffentlicht 21. Aug. 2026 · Aktualisiert 26. Aug. 2026
Voraussetzungen
Add the field dictionary with types, required flags, aliases, and allowed values. Define source identifiers, repeating-field layout, date, number, currency, and missing-value rules.
Skill-Dokument
Das vollständige SKILL.md, das dein Agent liest und befolgt.
Document field extraction
Purpose
Turn a heterogeneous document set into one row per source document and one column per declared field. Preserve the raw value, normalized value, source location, and extraction note so downstream users can audit a date, amount, identifier, party name, or line-item total without reopening the entire corpus.
Scope
Extract only the schema declared in the mapping document or run input. Handle repeated labels, tables, headers, footers, multi-page records, handwritten or unreadable values, and fields that are absent. Keep document-level fields separate from repeating line items and never merge two source documents because their names appear similar.
Excluded: classifying documents into business categories, deciding whether a value is legally valid, correcting a source record, and filling a blank from external knowledge.
Data basis
- The fixed document corpus, including file name or document identifier.
- Field dictionary with field_key, type, required flag, normalization rule, and allowed values.
- Optional mapping table for document type, page numbering, table headers, and source-system identifiers.
Result
Write an extraction sheet with document_id, field_key columns, raw value, normalized value, unit, page or section citation, confidence note, and missing or ambiguous status. Add a compact exceptions document for unmapped labels, duplicate candidates, unreadable areas, and schema mismatches.
Quality criteria
- Every source document has a row or a documented exclusion reason.
- Required fields are populated only when the source contains evidence; otherwise they say missing.
- Dates, numbers, currencies, and identifiers follow the field dictionary without losing raw text.
- Repeated fields retain all occurrences with an occurrence index or line-item row.
- Every populated value cites page, heading, table, row, or another reproducible location.
Instructions
Apply the field dictionary before interpreting labels. Prefer a value printed beside the matching label over a footer, filename, or inferred value. For conflicting occurrences, retain each candidate and mark the field ambiguous. Normalize decimal separators, date formats, and whitespace only under a declared rule. Never convert a blank, dash, or “N/A” into zero unless the mapping explicitly says so. Record a confidence note based on evidence quality, not on a guessed probability, and keep low-confidence values visible for review.
Adapt before use
- Add the exact field dictionary, data types, required flags, and allowed values.
- Map source identifiers, page conventions, and repeating table fields.
- Define normalization rules for dates, decimal separators, currencies, and identifiers.
- Set the vocabulary for missing, ambiguous, unreadable, and excluded values.
Verwandte Skills
- Document classification register
Classifies each document against a supplied taxonomy and records document type, business purpose, sensitivity, retention cue, owner, and rationale. Use for records inventories, contract libraries, information-governance reviews, and knowledge-base cleanup.
- Document comparison matrix
Compares two or more controlled documents clause by clause, recording additions, deletions, changed wording, changed values, and unresolved mappings with page or section citations. Use for policy revisions, contract redlines, procedure updates, and version-control reviews.
- Document translation and localization
Produces a faithful business translation in one or more target languages while preserving structure, terminology, figures, and formatting decisions. Use for contracts, policies, proposals, manuals, reports, and internal documents; keywords include translation, localization, multilingual, glossary, and terminology consistency.
- Patientenquittung für eine Versicherungs-Risikovoranfrage auswerten
Wertet eine Patientenquittung der gesetzlichen Krankenkasse aus und befüllt damit den Gesundheitsfragebogen eines Versicherungsmaklers; offene Punkte werden als Rückfrage an den Kunden markiert.