Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

ML Engineer

Approvalmax
CompanyApprovalmax
CategoryEngineering
Location
Remote
EmploymentNot stated
LevelNot stated
SalaryNot stated by the employer
Posted2 Jul 2026
Last verified30 Jul 2026
SourceEmployer career page (teamtailor)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
About ApprovalMax ApprovalMax is a fast-growing B2B SaaS company that helps businesses automate their approval workflows and financial controls. With a global team of over 150 people spanning the UK, Europe, North America, Australia, and South Africa, we build software that matters and we’re scaling quickly. The Role Our Capture product extracts structured data from hundreds of thousands of financial documents monthly - invoices, bills, POs - through an OCR pipeline that matches extracted fields against customer accounting systems. Your KPI is zero-touch rate: the percentage of documents where the system output requires zero manual correction. Your job is to move it up - systematically, measurably, and permanently. We’ve built the foundation, a validated accuracy measurement framework on full production data, a comprehensive error taxonomy of root causes, an error identification methodology, and the first shipped production fixes. You inherit the methodology and the backlog. We need a dedicated owner to execute and scale it. The work splits roughly 70% forensic data investigation / 30% ML engineering, shifting toward 50/50 as models go to production. Four error origins drive the roadmap: Entity matching (~50% of fixable errors). OCR extracts field values correctly, but the pipeline matches them to the wrong account, supplier, or tax code. Planned: embedding-based similarity search, recommender systems, consensus-based coding prediction - a standalone ML service the core pipeline calls. Pipeline logic (~25%). Our post-processing pipeline introduces errors through its own deterministic logic - tax treatment misclassification, rounding, spurious adjustment lines. Planned: forensic investigation per pattern, tracing data through processing steps, designing and validating rule-based fixes. OCR extraction (~25%). The OCR engine misreads the document - wrong currency, phantom line items, structural parsing failures. Planned: build an OCR correction layer - the right approach may be LLM with guardrails, an alternative OCR engine, a correction model from HuggingFace, or a combination. Freedom to choose; rigour required to validate. User overrides (~equal to the above combined, lower priority). Users change correct values for business reasons. Future: learn organisation/vendor correction patterns, build recommendation systems from historical data. Remote - applicants must be based in the UK, Serbia, or Moldova. Key Responsibilities: Accuracy Investigation & Measurement (~70% initially) Investigate why documents fail at population scale - query production datasets, compare multiple data representations per document, find statistical patterns that explain hundreds of failures at once. Balance population-level analysis with individual-document forensics where needed. Own and evolve the accuracy measurement framework. Every fix has an expected uplift, a measured uplift, and a post-deployment monitoring plan. Inherit and improve the error identification methodology. Two modes: LLM-assisted analysis for discovering new patterns across large document batches, and direct SQL investigation for patterns with clean statistical signatures. The methodology is proven and documented; you upgrade it with your own analytical instincts and DS expertise. Design fixes for pipeline logic errors by reading the C# codebase, understanding processing step sequence, and identifying root causes. Hand validated designs to C# engineers for production implementation. Verify measured results. ML Engineering & Model Development (~30% initially, growing) Build an embedding-based entity matching service: encode supplier/description signals into vector representations, evaluate retrieval quality against ground truth, iterate on ranking. Deploy as a Python service integrated with the core C# pipeline. Build an OCR correction layer to fix extraction errors before pipeline processing. E
HOUSE ADYou found the opening. Now track it.Tracker, radar and AI drafts in one place.erioun.com →