Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Senior Machine Learning Engineer, Synthetic Data & Document Understanding

ABBYY
CompanyABBYY
CategoryEngineering
LocationBangalore
RemoteHybrid
EmploymentNot stated
LevelSenior
SalaryNot stated by the employer
Posted21 May 2026
Last verified2 Aug 2026
SourceEmployer career page (greenhouse)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Join ABBYY and be part of a team that celebrates your unique work style. With flexible work options, a supportive team, and rewards that reflect your value, you can focus on what matters most – driving your growth, while fueling ours. Our commitment to respect, transparency, and simplicity means you can trust us to always choose to do the right thing. As a trusted partner for purpose-built AI and intelligent automation, we solve highly complex problems for our enterprise customers and put their information to work to transform the way they do business.  Over 10,000 customers trust ABBYY, including many Fortune 500 ones. You will work on further developing a portfolio already containing client names such as DHL, Johnson & Johnson, FDA, DMV, PwC, KeyBank, Spotify, and H&R BLOCK. About the Role   We are seeking a  Senior Machine Learning Engineer – Synthetic Data & Document Understanding  to own the synthetic data generation track within ABBYY’s  Document AI Data team .   This role focuses on building  generative pipelines that produce high-quality, diverse, and realistic synthetic training data  at scale. You will ensure synthetic data meaningfully improves downstream model performance by maintaining strong alignment with real-world document structures, formats, and statistical properties.   This is an ideal role for engineers who combine  deep generative modeling expertise with rigorous data quality evaluation and production engineering skills .   Key Responsibilities   Technical Development & Innovation   Design and implement pipelines that analyze real documents to inform  high-fidelity synthetic data generation   Build generative systems capable of producing documents across diverse  formats, layouts, and domains   Develop  evaluation frameworks  to ensure synthetic data maintains distributional fidelity and diversity   Research and apply  generative modeling techniques  suited for document AI training   Identify and mitigate quality issues to ensure synthetic data is effective for downstream model training   Partner with Modeling teams to measure the  impact of synthetic data on model performance   Project Ownership & Leadership   Own the synthetic data generation track  end-to-end , from architecture to quality validation   Drive architectural decisions balancing  quality, diversity, scale, and cost efficiency   Define and maintain  data quality metrics and generation dashboards   Collaborate closely with annotation teams to ensure compatibility with downstream pipelines   Contribute to roadmap planning alongside Principal-level leadership   Infrastructure & Scale   Build scalable pipelines capable of generating  millions of synthetic training examples   Implement  post-processing, filtering, and validation mechanisms  to remove low-quality outputs   Design cost-efficient workflows balancing  compute, quality, and throughput   Develop monitoring systems to detect  distribution shifts or quality degradation over time   Collaborate with Platform teams on  compute orchestration, storage, and scheduling   Qualifications   Education & Experience   MS or PhD in Computer Science, Engineering, Mathematics, or related field   5+ years of experience in  Machine Learning / AI , with focus on:    Generative models   Vision-Language Models (VLMs)   Synthetic data systems   Proven experience
HOUSE ADYou have the idea. We build it.