Marcus Feld
AI Evaluation Engineer
Berlin, Germany
Summary
Evaluation engineer who decides whether a model is good enough to ship. Built the harness gating every release at a 200-person AI company — 40+ benchmarks, offline and online. Caught two regressions that would have reached production and cut evaluation turnaround from four days to under three hours.
Experience
AI Evaluation Engineer · Halden AI
Jun 2023 – Present
- Own the evaluation harness gating every model release: 40+ automated benchmarks across capability, safety and regression.
- Cut evaluation turnaround from four days to under three hours by parallelising runs and caching deterministic scoring.
- Caught two capability regressions pre-release that aggregate metrics missed, by adding task-level breakdowns.
- Built the LLM-as-judge pipeline with human spot-checks and published agreement rates so teams know how far to trust it.
Machine Learning Engineer · Halden AI
Sep 2021 – May 2023
- Trained and fine-tuned retrieval and classification models, owning the data pipeline end to end.
- Built the first internal eval scripts — the work that became a dedicated evaluation function.
Education
MSc Computer Science
Technical University of Munich · 2016 – 2018
BSc Mathematics
University of Freiburg · 2013 – 2016
Certifications
Skills
Evaluation: Benchmark Design · LLM-as-Judge · Human Evaluation · Regression Testing · Error Analysis
ML Engineering: Python · PyTorch · Hugging Face · Fine-tuning · RAG
Infrastructure: CI/CD · Docker · Kubernetes · Weights & Biases · MLflow