How to Create an Automated Invoice Processing Agent Using Open-Source AI

An automated invoice processing agent does not need to begin as a large autonomous AI system. For a first useful version, think of an agent as a controlled software workflow that receives an invoice, calls the right document and AI tools, validates the result, decides whether a human must review it, and only then saves or exports approved data.

That distinction matters in finance. A language model can help turn messy invoice layouts into structured fields, but it should not be the only thing deciding whether an invoice is valid or safe to post. The beginner-friendly design in this guide keeps extraction probabilistic and the important accounting checks deterministic.

By the end, you will have a local prototype that accepts PDF or image invoices, extracts fields such as vendor, invoice number, dates, currency, totals, and line items, validates those fields against rules, routes questionable documents to review, stores an audit trail, and exposes the process through an API.

What do you need to understand before you start?

There are four terms worth knowing:

  • OCR (optical character recognition) turns text visible in an image or scan into machine-readable text.
  • Document parsing preserves more structure than plain OCR, such as reading order, tables, headings, and layout regions.
  • LLM (large language model) can map variable invoice text into a consistent JSON structure.
  • Human-in-the-loop means the system stops and asks a person to review cases that fail checks instead of pretending every result is correct.

As of September 2026, several actively maintained open-source projects cover these layers. Docling supports PDF, image, office-document, and other document formats; PaddleOCR provides OCR and document-understanding pipelines and is released under Apache 2.0; and the Tesseract project documents its current 5.x line as an open-source OCR engine in the official Tesseract manual.

For local model inference, llama.cpp's server documentation describes OpenAI-compatible endpoints and schema-constrained JSON output. That makes it useful for an extraction service because you can ask a locally hosted model to return a defined object rather than unconstrained prose.

One licensing warning is important: an open-source inference engine does not automatically make every model weight open source. Check the license for the exact model checkpoint you deploy, especially for commercial use.

Which stack should a beginner use?

Layer Practical choice Why it is here
Document conversion Docling Converts PDFs and images into a structured document representation and exportable text.
OCR fallback PaddleOCR or Tesseract Useful when the invoice is a scan or photo and text extraction needs OCR.
Local AI extraction llama.cpp + a suitable instruction model Maps variable invoice text to a stable JSON schema while keeping inference local.
Schema validation Pydantic Validates Python data against declared field types and constraints.
API FastAPI Accepts invoice uploads and returns processing results.
Storage SQLite for a prototype; PostgreSQL for a shared service Stores normalized data, review status, source-file identity, and audit events.
Orchestration Plain Python first; LangGraph optional later A deterministic function is easier to debug. Add a stateful agent framework when branching and human review become more complex.

Pydantic's official documentation explains that models inherit from BaseModel, validate incoming data, and can emit JSON Schema through model_json_schema(). See Pydantic model validation. If you later need durable, stateful branching, LangGraph's reference describes support for long-running stateful agents, persistence, and human-in-the-loop workflows.

Step 1: Set up the project and define a narrow first goal

Do not start with “automate accounts payable.” Start with one document class and one output contract. A good MVP is: English-language supplier invoices, PDF or image input, one legal entity, one currency set, and no automatic payment execution.

mkdir invoice-agent
cd invoice-agent
python -m venv .venv

# Linux/macOS
source .venv/bin/activate

pip install docling pydantic fastapi uvicorn httpx python-multipart

Install your OCR engine separately if you need one. Docling's current installation guide uses pip install docling; PaddleOCR's current installation documentation distinguishes its base OCR package from optional document parsing and information-extraction dependency groups. Use the Docling installation guide or the PaddleOCR 3.x installation guide rather than copying an old dependency list.

Example project setup panel showing an invoice-agent folder, Python virtual environment, and package installation commands
Step 1: Create a small isolated Python project first. Keep the first invoice type and output schema narrow enough to test thoroughly.

Step 2: Convert invoices into text and layout data

Born-digital PDFs may already contain usable text. Scanned PDFs and photos require OCR. Your parser should hide that difference from the rest of the pipeline: later stages should receive normalized document text plus any useful layout or table information.

A minimal Docling conversion looks like this:

from docling.document_converter import DocumentConverter

converter = DocumentConverter()

def parse_document(path: str) -> str:
    result = converter.convert(path)
    return result.document.export_to_markdown()

Docling's official quickstart uses the same convert-then-export pattern. For OCR-heavy collections, test PaddleOCR or Tesseract on your own scans. PaddleOCR's project documentation recommends its newer PP-StructureV3 family for document analysis, while Tesseract can be called from the command line or API for printed-text OCR.

Example invoice used to test OCR, showing vendor details, dates, line items, subtotal, tax, and total
Step 2: Begin with readable sample invoices that contain the fields you plan to extract, then add harder scans only after the basic pipeline works.

Step 3: Decide exactly what data the agent is allowed to produce

The extraction schema is the contract between AI and accounting logic. Avoid a giant “anything that might appear on an invoice” schema. Start with fields you can validate.

from datetime import date
from decimal import Decimal
from pydantic import BaseModel, Field, ConfigDict

class LineItem(BaseModel):
    description: str
    quantity: Decimal | None = None
    unit_price: Decimal | None = None
    amount: Decimal

class Invoice(BaseModel):
    model_config = ConfigDict(extra="forbid")

    vendor_name: str
    invoice_number: str
    invoice_date: date
    due_date: date | None = None
    currency: str = Field(min_length=3, max_length=3)
    subtotal: Decimal | None = None
    tax: Decimal | None = None
    total: Decimal
    line_items: list[LineItem] = []

extra="forbid" is useful in this context because unexpected LLM fields should not silently become part of your accounting record. Also note that Pydantic validation confirms structure and types; it does not prove that the extracted vendor or amount matches the invoice image.

Example invoice layout analysis with labeled vendor, address, invoice number, date, due date, line items, and total regions
Step 3: Define the target fields around business needs, then use layout and table information to preserve where those values came from.

Step 4: Use a local model to map document text into structured JSON

Run a suitable instruction model behind llama.cpp's HTTP server. The exact model size depends on your hardware and invoice complexity. Do not assume a larger model is automatically better; benchmark field accuracy, latency, and memory on your own invoice set.

llama.cpp currently supports schema-constrained JSON responses. You can derive that schema from Pydantic and send it with the extraction request:

import json
import httpx

def extract_invoice(document_text: str) -> Invoice:
    schema = Invoice.model_json_schema()

    prompt = f"""
Extract one supplier invoice from the text below.
Do not invent values. If an optional field is absent, use null.
Currency must be a three-letter code.
Return only data that fits the required schema.

DOCUMENT:
{document_text}
"""

    payload = {
        "model": "local-model",
        "messages": [{"role": "user", "content": prompt}],
        "temperature": 0,
        "response_format": {
            "type": "json_schema",
            "schema": schema
        }
    }

    r = httpx.post(
        "http://127.0.0.1:8080/v1/chat/completions",
        json=payload,
        timeout=120
    )
    r.raise_for_status()

    content = r.json()["choices"][0]["message"]["content"]
    return Invoice.model_validate_json(content)

Schema-constrained generation helps with valid structure, but it does not guarantee semantic truth. A perfectly valid JSON object can still contain the wrong total. That is why the next step is mandatory.

Example structured JSON output containing vendor, invoice number, dates, currency, totals, and line items
Step 4: Ask the local model for a strict structured object, not free-form prose, so downstream validation receives predictable fields.

Step 5: Add deterministic accounting checks and review rules

Do not ask the LLM “Are you confident?” and trust the answer. Self-reported model confidence is not a calibrated control. Build checks that can be recomputed from the extracted data and your systems.

Useful first rules include:

  • required fields are present;
  • invoice number is not already recorded for the same vendor;
  • currency is allowed for that vendor or entity;
  • line-item amounts approximately sum to subtotal;
  • subtotal plus tax approximately equals total;
  • invoice date and due date are plausible;
  • vendor exists in your vendor master or is routed for onboarding;
  • purchase-order references match when PO matching is required.
from decimal import Decimal

def validate_business_rules(inv: Invoice) -> list[str]:
    errors = []

    if inv.subtotal is not None and inv.tax is not None:
        expected = inv.subtotal + inv.tax
        if abs(expected - inv.total) > Decimal("0.02"):
            errors.append("subtotal_plus_tax_does_not_match_total")

    line_sum = sum((x.amount for x in inv.line_items), Decimal("0"))
    if inv.subtotal is not None and abs(line_sum - inv.subtotal) > Decimal("0.02"):
        errors.append("line_items_do_not_match_subtotal")

    if inv.due_date and inv.due_date < inv.invoice_date:
        errors.append("due_date_before_invoice_date")

    return errors

Route every failed check to a review queue. For a first deployment, it is reasonable to route all invoices to review while you measure extraction quality. Automation thresholds should come from evidence collected on your own labeled invoices, not from a generic percentage copied from another system.

Example invoice validation panel with checks for required fields, total arithmetic, due date, currency, and numeric amounts
Step 5: Treat arithmetic, duplicate detection, vendor matching, and date rules as deterministic controls; route exceptions to a person.

Step 6: Wrap the processor in a small API

An API lets an upload form, email worker, shared folder watcher, or accounting integration call the same processing function. FastAPI's official file-upload guide recommends UploadFile for uploaded files and notes that it uses a spooled file, which is better suited than loading large files entirely into memory.

from pathlib import Path
from tempfile import NamedTemporaryFile
from fastapi import FastAPI, UploadFile, HTTPException

app = FastAPI()

@app.post("/invoices")
async def process_upload(file: UploadFile):
    if file.content_type not in {
        "application/pdf",
        "image/png",
        "image/jpeg"
    }:
        raise HTTPException(415, "Unsupported file type")

    suffix = Path(file.filename or "").suffix

    with NamedTemporaryFile(suffix=suffix, delete=False) as tmp:
        tmp.write(await file.read())
        path = tmp.name

    result = process_invoice(path)
    return result

See the current FastAPI request-file documentation for the supported upload patterns. In production, also enforce file-size limits, malware scanning where appropriate, authenticated callers, request IDs, and secure temporary-file cleanup.

Example FastAPI code panel for a POST invoices endpoint and a successful HTTP response
Step 6: Expose one processing endpoint after the local pipeline works, so multiple input channels can reuse the same validation path.

Step 7: Store the original evidence, normalized result, and audit events

Never store only the final JSON. Keep enough evidence to reconstruct what happened: original file identity, hash, parser version, model identifier, prompt version, extracted JSON, validation results, review decision, timestamps, and the user or service that approved a change.

A simple prototype can use SQLite. A multi-user service can use PostgreSQL with uniqueness and foreign-key constraints. PostgreSQL's current documentation covers table constraints, which are useful for protecting keys and relationships independently of the AI pipeline.

CREATE TABLE invoices (
    id BIGSERIAL PRIMARY KEY,
    vendor_id BIGINT NOT NULL,
    invoice_number TEXT NOT NULL,
    invoice_date DATE NOT NULL,
    currency CHAR(3) NOT NULL,
    total NUMERIC(18,2) NOT NULL,
    status TEXT NOT NULL,
    source_sha256 TEXT NOT NULL,
    extracted_json JSONB NOT NULL,
    created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
    UNIQUE (vendor_id, invoice_number)
);

The source hash gives you a second way to detect accidental resubmission. Keep file storage permissions tighter than general application logs because invoices often contain addresses, bank details, tax IDs, and other business-sensitive data.

Example database tables for invoices and source invoice files with status and timestamps
Step 7: Persist both normalized invoice data and source-file metadata so approvals, duplicates, and later corrections remain auditable.

Step 8: Test with real invoice variation before automating downstream actions

A prototype that works on five clean PDFs is not production-ready. Build a labeled evaluation set that represents the inputs you actually receive: different vendors, scans, rotated pages, small fonts, multi-page invoices, discounts, negative lines, taxes, shipping charges, multiple currencies, missing purchase orders, and duplicate submissions.

Track field-level metrics separately. “Invoice accuracy” can hide the fact that vendor name is easy while line-item quantity is unreliable. At minimum, measure:

  • exact-match accuracy for invoice number, currency, and dates;
  • numeric tolerance accuracy for subtotal, tax, and total;
  • line-item precision and recall if line items matter downstream;
  • percentage routed to human review;
  • false auto-approval rate;
  • processing latency and failure rate;
  • duplicate-detection precision.

Store failures as test cases. Every time you fix a parser rule, prompt, schema, or model, rerun the whole evaluation set. This prevents one vendor's improvement from quietly breaking another vendor's layout.

Example invoice processing review dashboard with processed, needs review, and failed counts plus recent invoice statuses
Step 8: Measure the review queue and failure cases on representative invoices before allowing the agent to automate downstream accounting actions.

How should the agent decide what happens next?

Keep the state machine explicit. The “agent” can be a plain Python function at first:

def process_invoice(path: str) -> dict:
    text = parse_document(path)
    invoice = extract_invoice(text)
    errors = validate_business_rules(invoice)

    if errors:
        status = "needs_review"
    else:
        status = "validated"

    record_id = save_invoice(
        source_path=path,
        invoice=invoice,
        validation_errors=errors,
        status=status,
    )

    return {
        "id": record_id,
        "status": status,
        "invoice": invoice.model_dump(mode="json"),
        "validation_errors": errors,
    }

This is already agent-like because it coordinates tools and chooses a branch. You do not need an LLM to choose every next step. Add LangGraph or another orchestration framework when you need persisted multi-stage reviews, retries, asynchronous callbacks, or several specialized tools whose execution order changes by case.

What mistakes should beginners avoid?

Letting the LLM be the accounting control

The model should extract and normalize; code and trusted system data should verify. Arithmetic, duplicate checks, vendor status, PO matching, and approval limits belong in deterministic logic.

Discarding the source document after extraction

Keep the original invoice or a controlled reference to it. Reviewers need evidence, and you need a way to reproduce failures after model or parser upgrades.

Using one prompt without a version

Store a prompt version and model identifier with each processed invoice. Otherwise, you cannot explain why two invoices processed months apart behaved differently.

Automatically posting every schema-valid result

JSON validity is not invoice validity. Start with human review, then automate only low-risk cases after your evaluation demonstrates acceptable performance.

Ignoring model and data licenses

Libraries and model weights can have different licenses. PaddleOCR and Tesseract publish open-source licenses, while the license of the model you run through llama.cpp depends on that specific checkpoint. Confirm it before deployment.

Building the UI before proving extraction quality

A polished dashboard cannot compensate for incorrect amounts. Test document parsing, schema extraction, and validation first; add a review interface after you know what reviewers actually need to see.

When should you use OCR-only, layout models, or an LLM?

Approach Best fit Main limitation
OCR + rules Small number of highly stable vendor templates Rules become brittle as layouts vary.
Document layout model Tables, forms, and fields where position matters May need task-specific training or post-processing.
OCR/document parser + local LLM Many invoice formats with a common target schema Must be constrained and validated; inference costs more compute.
Vision-language model Complex documents where text and visual layout are tightly coupled Higher hardware requirements and additional evaluation complexity.

Hugging Face's current LayoutLMv3 documentation describes a document-AI model that combines text and visual layout information. PaddleOCR's PP-StructureV3 documentation likewise focuses on layout, tables, and structured document analysis. These are options when plain text plus an LLM is not enough.

What is a sensible path from prototype to production?

Move in stages. First, run locally on a folder of labeled invoices. Second, add the API and persistent storage. Third, introduce a human review queue. Fourth, integrate read-only lookups such as vendor master or purchase orders. Fifth, allow controlled export to an accounting staging area. Only after monitoring proves the controls should you consider automatic posting for tightly defined low-risk cases.

At every stage, preserve three things: evidence (the original document and extracted source), determinism (business rules that can be re-run), and traceability (which parser, model, prompt, and reviewer produced the final record).

Final checklist

  • The invoice parser works on both born-digital PDFs and the scans you actually receive.
  • The extraction output is constrained to a versioned Pydantic schema.
  • The model is not allowed to invent missing values.
  • Totals, dates, duplicates, vendor identity, and PO rules are validated outside the LLM.
  • Failed rules produce a human-review state instead of silent correction.
  • The original document, source hash, model, prompt version, and validation result are logged.
  • The API restricts file type and size and is authenticated before production use.
  • Your test set contains real layout variation and edge cases.
  • Model and library licenses have been checked for your intended deployment.
  • Downstream accounting writes are introduced gradually and remain auditable.

If those controls are in place, you have more than an OCR demo: you have the foundation of a dependable invoice processing agent. The open-source components can change over time, but the architecture remains durable—parse evidence, extract into a schema, validate with code, route uncertainty to people, and record every decision.

Leave a Comment

How to Create an Automated Invoice Processing Agent Using Open-Source AI

How to Create an Automated Invoice Processing Agent Using Open-Source AI

Build a practical open-source invoice processing agent with document parsing, local LLM extraction, validation, human review, an API, storage, and testing.

Why Your AutoGen Agent Gets Stuck in Infinite Loops—and How to Fix It

Why Your AutoGen Agent Gets Stuck in Infinite Loops—and How to Fix It

Diagnose AutoGen infinite loops by checking termination rules, speaker selection, tool retries, handoffs, state, and traces, with practical fixes for AgentChat.

Printable Daily Time Blocking Template PDF for WFH Professionals

Printable Daily Time Blocking Template PDF for WFH Professionals

Use a printable daily time blocking template for remote work, with focus blocks, meetings, breaks, buffers, and a shutdown routine that fits one page.

Best Free Invoicing Apps for Small Service Businesses in the US (2026)

Best Free Invoicing Apps for Small Service Businesses in the US (2026)

Compare the best free invoicing apps for US service businesses in 2026, including Zoho Invoice, Wave, Square, PayPal, and Invoice Ninja.

AI Agent Access Checklist for Small Businesses Before Fall Sales

AI Agent Access Checklist for Small Businesses Before Fall Sales

Before fall promotions, limit what an AI agent can read or change. Use this small-business checklist for permissions, customer data, approvals, testing, and offboarding.

Siri AI in iOS 27: What It Can Do and When It Still Asks You to Confirm

Siri AI in iOS 27: What It Can Do and When It Still Asks You to Confirm

See what Siri AI can find, draft, and do across supported iOS 27 apps, when it may ask for approval, and which settings and availability limits to check.

Preparing an AI Agent Demo for OpenAI DevDay 2026 Without Customer Data

Preparing an AI Agent Demo for OpenAI DevDay 2026 Without Customer Data

Build a DevDay-ready agent demo with fictional fixtures, limited tools, inspected traces, and a full rehearsal—without relying on live customer records.

Best Low-VRAM Settings for Running Llama 3 Locally on Mid-Range Laptops

Best Low-VRAM Settings for Running Llama 3 Locally on Mid-Range Laptops

Tune Llama 3 8B for 4–8 GB VRAM laptops with practical quantization, context, GPU offload, and batch settings that balance memory, speed, and response quality.

How to Fix a Word Document That Opens as Read-Only on Mac

How to Fix a Word Document That Opens as Read-Only on Mac

Fix Word documents that open read-only on Mac by checking Office updates, file permissions, cloud access, document restrictions, and shared-file locks.

How to Keep Character Consistency Across Multiple Runway Gen-3 Shots: A 2026 Workflow

How to Keep Character Consistency Across Multiple Runway Gen-3 Shots: A 2026 Workflow

Runway Gen-3 is retired, but its character-consistency problem remains. Use references, character plates, disciplined shot design, and image-to-video workflows to reduce drift.