Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

Principal Applied Scientist | LLM Benchmarking | Berlin / Munich / Remote in Germany

Xcede
CompanyXcede
CategoryScience & Research
LocationBerlin
RemoteRemote (inferred)
EmploymentNot stated
LevelLead
SalaryNot stated by the employer
First seen2 Aug 2026 (the employer did not state a posting date)
Last verified9 Aug 2026
SourceThe employer's own careers page (company_site)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
Principal Applied Scientist | LLM Benchmarking | Berlin / Munich / Remote in Germany Confidential search for a fast-growing, Series C agentic AI company building conversational AI for global enterprise brands. This isn't a research seat and it isn't a data analytics role. You'd own prompt optimisation and LLM benchmarking end to end: building the evaluation framework that decides which models the company adopts, comparing quality against latency across real production use cases, and generalising that framework so it can eventually judge any model, not just the ones already wired into the product. Six months from now, the ambition is for this to be a benchmark other companies reference. What you'll do • Improve and optimise prompts for real production use cases and agent workflows, systematically evaluating performance across different prompting strategies and models. • Build and maintain benchmarking scenarios that evaluate models (Gemini, GPT-class systems, and whatever comes next) against task performance, end-to-end system integration, latency, and quality. • Assess new model releases as they land, validate them against our performance and latency requirements, and give data-driven recommendations on whether we adopt them. • Design and evolve a generalised evaluation framework, starting from our current benchmarking tooling and gradually decoupling it from our core agent system so it can assess any model, including ones we haven't integrated yet. • Work closely with our agent and platform engineering teams to turn findings into production changes. • Track quality and latency trends across model versions over time. What you'll need: • Strong Python • Hands on experience with LLMs and prompt engineering • A real understanding of benchmarking and evaluation methodology • Systems thinking, you'll be building tooling, not running one off notebooks Research background is a plus, so is experience comparing models at scale. Not a fit if you're a pure analyst or a theoretical researcher who hasn't shipped to production. Berlin or Munich preferred, remote within Germany considered.