Job Opportunities API

The Public Ledger of Openings

← Back to the ledger

PhD Research Internship – Robotics Engineer (VLM / VLA Models)

Sensmore
CompanySensmore
CategoryEngineering
LocationBerlin / Potsdam
RemoteOn-site (inferred)
EmploymentNot stated
LevelIntern
SalaryNot stated by the employer
Posted23 Apr 2026
Last verified11 Aug 2026
SourceEmployer ATS (ashby)
Applications are handled by the employer, not by us.Apply on the employer's site →
Description
sensmore automates the world's largest machines with unprecedented intelligence. Our proprietary Physical AI enables heavy machines such as wheel loaders to instantly adapt to dynamic environments and execute new tasks without prior training. We integrate cutting-edge robotics into a platform powering intelligence and automation products - transforming productivity and safety for customers in mining, construction, and adjacent industries today. Join us and play a pivotal role in transforming the automation landscape in heavy industries. ROLE OVERVIEW We are seeking a highly motivated PhD candidate to join our team as a Research Intern specializing in General Purpose AI, with a focus on Vision-Language Models and Vision-Language-Action systems. This role sits at the frontier of industrial robotics: developing scalable, general-purpose VLA systems that enable robots to perceive, reason, and act autonomously in complex heavy-industry environments. You will contribute to bridging multi-modal perception (e.g., video, radar, lidar) with robust real-world execution, while advancing state-of-the-art methods in embodied AI. Beyond engineering, this position has a strong research component, with opportunities to contribute to novel methods, publish findings, and shape the future of industrial autonomy. KEY RESPONSIBILITIES Depending on your expertise and project priorities, you will: Research & Method Development - Design and develop novel approaches for Vision-Language-Action systems in real-world industrial settings - Explore scalable architectures for multi-modal reasoning and action generation - Contribute to advancing state-of-the-art methods in embodied AI and robotic autonomy Multi-Modal Learning & Data Systems - Lead the design and analysis of large-scale multi-modal datasets (video, radar, lidar, sensor fusion) - Develop self-supervised or weakly supervised dataset generation pipelines for VLA training - Investigate data-centric approaches to improve robustness and generalization Model Development & Optimization - Build, adapt, and extend cutting-edge GenAI models (e.g., VLMs, VLA frameworks) - Apply advanced fine-tuning strategies (e.g., parameter-efficient tuning, alignment methods) - Explore prompt optimization, reasoning augmentation, and action grounding techniques Training, Evaluation & Benchmarking - Design rigorous evaluation protocols for embodied AI systems in industrial contexts - Run large-scale experiments, analyze performance, and iterate systematically - Benchmark models against state-of-the-art approaches and internal baselines Deployment & Systems Integration - Collaborate with engineering teams to transition research prototypes into production-ready systems - Optimize models for real-time inference, robustness, and safety in heavy-industry environments Scientific Contribution - Document findings and contribute to research publications, technical reports, or patents - Present results internally and potentially at leading conferences REQUIRED QUALIFICATIONS - Current enrollment in a PhD program in Robotics, Computer Science, Machine Learning, Electrical Engineering, or a related field - Strong programming skills in Python and deep learning frameworks (e.g., PyTorch) - Solid understanding of machine learning, deep learning, and multi-modal models - Proven ability to conduct independent research and drive projects from idea to results - Strong analytical thinking and problem-solving skills PREFERRED SKILLS & EXPERIENCE - Experience with Vision-Language Models, embodied AI, or robotics learning systems - Familiarity with modern GenAI tooling (e.g., Hugging Face ecosystem, Gemini, Unsloth, or similar) - Experience with multi-modal data (vision + sensor fusion) - Background in robotics, control systems, or real-world deployment - Track record of research output (publications, preprints, or significant research projects) - E