Senior Data Engineer Python/GCP (x/f/m)
Doctolib
| Company | Doctolib |
| Category | Engineering |
| Location | Paris |
| Remote | On-site (inferred) |
| Employment | Not stated |
| Level | Senior |
| Salary | Not stated by the employer |
| Posted | 5 Jan 2026 |
| Last verified | 30 Jul 2026 |
| Source | Employer career page (greenhouse) |
Description
Your Impact
We are looking for a Senior Data Engineer to join the AI Team working on our AI Medical Companion .
Your mission will be to build and optimize the data foundations that power safe, scalable, and impactful AI models. You will work on data infrastructure for LLM, VLM, and RAG-based systems , ensuring our engineers and data scientists can train, evaluate, and deploy AI models efficiently on high-quality, well-structured, and compliant data. Your work will directly support health professionals in delivering better care while improving their work-life balance, ultimately impacting 80 million patients and 400,000 healthcare professionals across Europe.
Working in the tech team at Doctolib means building innovative products and features to improve the daily lives of care teams and patients.
What you'll do
Your responsibilities include but are not limited to:
Design, build, and maintain scalable data pipelines on Google Cloud Platform (GCP) for AI and machine learning use cases
Implement data ingestion and transformation frameworks that power Retrieval systems and training datasets for LLMs and multimodal models
Architect and manage NoSQL and Vector Databases to store and retrieve embeddings, documents, and model inputs efficiently
Collaborate with ML and platform teams to define data schemas, partitioning strategies, and governance rules that ensure privacy, scalability, and reliability
Integrate unstructured and structured data sources (text, speech, image, documents, metadata) into unified data models ready for AI consumption
Optimize performance and cost of data pipelines using GCP native services (BigQuery, Dataflow, Pub/Sub, Cloud Storage, Vertex AI)
Contribute to data quality and lineage frameworks, ensuring AI models are trained on validated, auditable, and compliant datasets
Continuously evaluate and improve our data stack to accelerate AI experimentation and deployment
Who you are
Before you read on: if you don't have the exact profile described below, but you feel this job description matches your skill set, we still encourage you to apply.
You'll be a great fit if you:
You have 5+ years of experience in Data Engineering, ideally supporting AI or ML workloads
You have strong experience with the GCP data ecosystem and proficiency in Python and SQL
You have deep understanding of NoSQL systems (e.g., MongoDB) and vector databases (e.g., FAISS, Vector Search)
You have experience designing data architectures for RAG, embeddings, or model training pipelines
You have knowledge of data governance, security, and compliance for sensitive or regulated data
You are fluent in English
It would be fantastic if you:
You hold a Master's or Ph.D. degree in Computer Science, Data Engineering, or a related field
You have familiarity with W&B / MLflow / Braintrust / DVC for experiment tracking and dataset versioning
You have experience with containerized environments (Docker, Kubernetes) and CI/CD for data workflows
Life at Doctolib Tech
Our solutions are built on a single fully cloud-native platform that supports web and mobile app interfaces, multiple languages, and is adapted to country and healthcare specialty requirements.
Our stack is composed of Rails, TypeScript, Java, Python, Kotlin, Swift, and React Native.
We leverage AI ethically across our products to empower patients and health professionals. Discover our AI vision here .
Want to learn more about our tech culture and environment? Visit the Doctolib Tech site .
What we offer
Free comprehensive health insurance (basic package) for you and your children
25 days of paid vacation per year, plus up to 14 days of RTT
Free mental health and coaching services through our partner Moka.care
Work f
995,367 openings. Erioun finds yours.Scored against your own profile, every hour.Try the radar →