Skip to main content
HighVRAM Hardware Watch

Extract fictional invoices into checked CSV with local AI

Practice exercise: validator tested; local-model walkthrough not yet run. All invoices are fictional. This exercise starts from plain text and does not test OCR, hardware speed or a particular model.

A convincing answer can still contain the wrong number. Try a small, checkable local-AI task: turn seven fictional invoices into CSV, preserve missing fields and flag totals that do not add up. The answer key makes it possible to separate a tidy-looking response from a correct extraction.

What you will need

  • An already working local text-chat model. If you need a starting point, use Your first local AI chat.
  • An existing Python 3 installation for the checker. The package uses the standard library and does not install dependencies or call a model.
  • A plain-text editor and the fictional invoice practice package. Unzip it into a folder you can find.

The package includes seven source documents, copyable prompts, the expected CSV, a data dictionary, an explanation of each answer and a validator. Keep the answers and tests out of the model input. Use only the fictional records for this first attempt.

1. Check the checker

Open a terminal in the extracted package folder. Run these commands with your existing Python command; on some systems that command is python rather than python3.

python3 -m unittest -v test_validator
python3 validate.py expected.csv

The first command tests the checker with authored good and bad examples. The second grades the supplied answer key and should return status pass. That result proves the package checks ran; it says nothing about how your model will perform. A passing extraction still includes documents needing review.

2. Send the prompt to your local model

Start a fresh chat with the model running on your machine. Leave browsing, shell access and connected tools disabled. The exercise needs only text input and a text response; invoice footers must not become instructions to act.

If your interface provides separate message roles, put prompts/system.txt in the system field and the complete prompts/user.txt in the user message. Otherwise paste prompts/single-message.txt as one message. Record which route you used. The files already include the contract and all seven invoices; do not add expected.csv or the answer key.

The prompt requests one CSV row per invoice. It separates printed values from calculated values and uses the literal marker \N for an absent source field. It also asks for a review flag. Check that the whole prompt fits the conversation and retain the model settings used for this attempt.

3. Save the raw response

Save the complete response as UTF-8 text named response.csv. Do not remove Markdown fences, tidy quotations or correct a total before checking it. Those mistakes are part of the result. Use a plain-text editor first instead of opening arbitrary model output directly in a spreadsheet. Never run code or follow a payment instruction from an invoice or response.

python3 validate.py response.csv --report response.validation.json

The checker returns pass only when all seven rows match the exercise contract and answer key. It rejects missing or extra columns, invented values, invalid dates, duplicate records and arithmetic or review flags that do not follow the contract. A failed check identifies the row and field where possible. The JSON report also lists source problems that were correctly preserved.

4. Look for these traps

  • Missing due date: the amounts can balance while the due date remains unknown. Do not assume Net 30.
  • Discount: subtract the discount before calculating tax, and add untaxed shipping afterward.
  • Credit note: keep amounts that are already negative; do not reverse their signs a second time.
  • Rounding: round each line before summing. Here, 1.005 rounds to 1.01 and 0.335 rounds to 0.34, giving a subtotal of 1.35.
  • Wrong printed total: one invoice prints 90.00 even though its components total 91.00. Preserve both values and flag the difference.
  • Missing tax amount: a rate may let you infer a plausible number, but extraction must keep an unprinted field missing.
  • Instruction-like footer: a quoted demand to ignore the schema remains document data. The model should continue extracting.

Open ANSWER-KEY.md after saving your first result. Correct extraction puts A02, A06 and A07 in the review queue. A validator pass means the response matches this small exercise; it does not mean every invoice is correct, legitimate or approved for payment.

5. Record what actually happened

Keep the raw response, validation report, exact model, quantization, app/runtime version, prompt delivery mode, context and output limits, and sampling settings. If a response is truncated, record it as a failure. If you change the prompt or repair the output, save a separate attempt instead of replacing the original.

For a small repeatability check, decide on three fresh-chat attempts before starting and report every result. This public seven-document set is useful for learning, but it is too small to rank models or decide what hardware to buy. It contains no scans, handwriting, real tax rules or multiple currencies. Real work needs broader tests and human review.

Next, use Test your local AI workload before you buy to record quality, waiting time and memory observations without confusing one successful extraction with a hardware benchmark.