MUHAMMAD TALHA Book a call

Document extraction

Your invoices land in QuickBooks or Xero without anyone retyping them.

Supplier invoices, bills of lading, receipts and contracts, read into clean, structured data. Fields the model isn't sure about go to a review queue, not into your books.

Attach the files or send a shared folder link. Worst scans first: they show the real accuracy.

How it works

Start with 20 of your files, not a sales deck.

  1. You send 20 real files. The messy ones: scans, phone photos, multi-page PDFs, odd vendor layouts.
  2. I get extraction running on them and check every value against the page it came from. A value that can't be found in the source is flagged, never guessed.
  3. You get the per-field accuracy report in 1 to 2 weeks: for each field (vendor, invoice number, dates, line items, tax, totals), how often it was right and which files it struggled with.
  4. Then the hand-off. Approved data goes into QuickBooks Online, Xero, your ERP or a sheet, and low-confidence fields wait for a person to check them.

Fixed price for the pilot, quoted after I see your files.

Proof

I've built this twice.

ShipMatch

2026 · product build · public repo

Accounts payable for import shipments. It reads invoices, bills of lading and customs entries from email, groups each into its shipment, runs 38 checks for overcharges, duplicates and duty math, and posts approved bills to QuickBooks Online or Xero. AI only reads the documents; matching and money math are plain, tested code. 1,054 automated tests; 6 of 6 planted errors caught with zero false alarms on its labelled benchmark.

Public repo on GitHub ↗

FirstGlance

2025–present · CTO · sole engineer

A Claude pipeline in production that pulls 30+ validated fields out of messy pitch-deck PDFs. Every result is checked against a schema and retried up to 3 times when it fails, so bad extractions are caught before they reach the database. Handles 20MB+ uploads. 100+ B2B companies onboarded.

firstglance.io ↗

Straight answer: ShipMatch's numbers come from its own labelled benchmark, not your documents, and scanned paper adds an OCR step. That's why the pilot runs on your worst files first: you see real accuracy on your data before you commit to anything bigger.

Next step

Send 20 files, get the per-field accuracy report.

Email them over with one line on where the data should end up. I'll reply with what I see and a fixed price for the pilot. Or book 30 minutes first.

Also: Lovable or Cursor app to production MCP and AI integration About me