R&D Engineering Manager, AI Evaluation
Apply or review the details on the original posting.
Apply for this roleOriginal posting on LinkedIn
About Us
Resaro was founded on the belief that AI will change the world in ways we cannot even imagine - but every new technology needs safeguards to advance. We are an independent, third-party AI assurance company: we build the software and run the evaluations that let enterprises and public-sector bodies deploy AI they can actually trust. Our work spans computer vision, generative AI and LLMs, vision-language models, and increasingly agentic and autonomous systems, for clients across government, defence, and commercial sectors.
Our product, the Approved Intelligence Platform (AIP), is where this becomes real software: customers upload datasets, register the AI systems they want tested, run rigorous evaluations, and produce defensible evidence and reports.
We're hiring a hands-on engineering leader for our AI evaluation team - someone whose primary focus is growing and managing a high-performing team of engineers and scientists, while maintaining the technical depth to guide their architecture and system design. You'll drive AI/ML engineering excellence and coordinate across teams to align technical decisions and resolve dependencies. You'll partner closely with the Product team.
This role can be based either in Singapore or Munich.
What You'll Do
Own the evaluation science, decide how we measure AI systems (accuracy, reliability, robustness, data quality) and turn research into methods that hold up under customer and regulatory scrutiny.
Lead the technical direction of AIP's evaluation engine across verticals: computer vision, LLM/RAG, VLM, and emerging areas like agentic and embodied AI systems.
Shape where the capability goes next, understand the systems we test and the customers who rely on the evidence, and bring a strong point of view to Product on which evaluation capabilities will matter 6-12 months out.
Lead through technical context. Your primary output is the success of your team. You set the bar for engineering quality through rigorous system design, strategic code reviews, and pairing - stepping into the codebase to unblock the team and guide architecture.
Own conceptual integrity as the system grows. Make and document the architecture decisions that matter, keep them coherent, and prevent uncontrolled coupling on key hotspots.
Make people management your top priority. Own 1:1s, performance, and career development for a multidisciplinary team of 8–10 engineers and scientists across Singapore and Europe. Partner on hiring to raise the team's bench strength, and create the shared context and ways of working that let engineers, AI engineers, and governance analysts solve problems together without constant top-down orchestration - building a culture of high trust and high output.
Own delivery and partner across teams. Translate the roadmap into executable plans, sequence the work, and set the standard for what reaches customers. This role reports to the CTO and works closely with the Product team to communicate load and dependencies, and to manage scoping and resourcing to ensure consistent execution against the product roadmap.
What We're Looking For
7+ years of professional software engineering experience shipping and operating production systems, including time as a tech lead and/or engineering manager.
Working depth in ML, data science, or statistics - enough to design and defend an evaluation metric and guide applied R&D, not just implement someone else's spec. You don't need to be a research scientist; you do need to hold your own with the ones on your team.
Proficiency in Python for backend services and data pipelines.
Demonstrated technical leadership of an engineering team - owning architecture decisions and setting engineering standards.
Direct people-management experience (or clear, evidenced readiness for it).
Strong architecture judgement - experience managing coupling, leading migrations, and keeping a growing system coherent through ADRs and dependency hygiene.
A leader's mindset with a builder's background - you find your deep satisfaction in growing people, scaling a team, and ensuring conceptual integrity.
Clear written and verbal communication, and comfort being measured against concrete quarterly outcomes.
Nice To Have
Experience building or integrating data-quality tooling, or evaluation/testing/assurance capability, in one or more AI domains: LLM/RAG, agentic systems, computer vision.
Ability to lead across a typed frontend stack (TypeScript/React).
Data-intensive pipelines with columnar/lakehouse formats (Parquet/Iceberg) and DuckDB or similar.
Container-based or serverless execution frameworks (Nuclio or comparable function/orchestration systems).
Kubernetes and Helm for production workloads.
GPU/CUDA infrastructure for model inference at scale.
Leading a platform migration with live, enterprise customers where evidence, auditability, and back-compatibility matter.
Published or applied research in model evaluation, benchmarking, or AI safety/assurance.
Our Hiring Blueprint
We hire to a high, transparent bar. We look for:
Production-grade code. You ship software that is correct, tested, observable, and maintainable by others
Communication and handover. You write clearly, document decisions, and leave work in a state another engineer can pick up. As a lead, you make your team's context legible.
Independent operation. You can take an ambiguous problem, scope it, decompose it, sequence it, and drive it to a shipped outcome without close supervision - and you help your team do the same.
T-shaped profile. Deep in your core domain (AI/backend/systems/etc.) and broad enough to operate steadily across the stack and the ML-evaluation domain
Resaro is an Equal Opportunity Employer. We respect each individual and support the diverse cultures, perspectives, skills and experiences within our teams.