Don’t get fooled by dark factory vendors — live session, Sept 29. Save my seat

Evaluation Lead – AI Observability

Full-Time
, Remote

Are you an LLM evaluation specialist eager to help engineering teams build AI products they can measure, trust, and continuously improve?

Join Camplight, where your expertise will help establish practical, credible evaluation standards for the next generation of AI applications.

What you’ll be working on?

We are building an AI Observability platform for teams developing LLM-based applications and agents. The platform helps teams understand how their systems behave and supports them in running meaningful evaluations.

Teams can often collect traces and build prototypes, but they need deeper expertise to answer the questions that decide whether an AI feature is ready: What should we measure? What does a credible dataset look like? When should we trust an LLM-as-a-judge? Is an apparent improvement real, or simply noise?

This initiative will help product teams make evaluation a repeatable engineering practice rather than an ad-hoc activity.

Your Role

Your role will involve becoming the evaluation expert within an AI Observability Centre of Excellence.

You will create practical guidance, reusable patterns, templates, worked examples, and reference implementations for evaluating LLM applications and agents. You will advise product teams on evaluation design, including task success, faithfulness and grounding, retrieval quality, tool and action correctness, safety, refusal behaviour, and regression testing against known failures.

You will guide teams in building golden datasets, defining annotation guidelines, choosing samples, calibrating LLM-as-a-judge approaches against human judgement, and interpreting variance and statistical confidence.

Product teams remain accountable for their own evaluation outcomes. You are not a release gate and will not run every team’s evaluations for them. Your role is to enable strong, credible evaluation practice at scale.

About Camplight

We build self-organizing technical teams, offer software development services, and work with businesses and entrepreneurs to create new products.

With over 300 successful software projects, some ongoing for over 8 years, we strive for long-term success for our partners.

By following the principles of self-management and organizing as a cooperative, we achieve 95% satisfaction among them.

We seek the best talent to join us and value transparency, collaboration, trust, responsibility, and innovation.

When joining Camplight, you can become a co-owner of the cooperative, allowing you to steer the business and share in the rewards of our collective success.

What are we looking for?

  • Ownership mindset: We want individuals who care about the quality of decisions that evaluation data supports. You do not treat a score as truth without understanding how it was produced.
  • Technical expertise: You can design credible evaluation approaches, understand uncertainty, and distinguish meaningful change from measurement noise.
  • Practical enablement: You can make complex evaluation concepts useful for busy product teams through templates, examples, and focused consultation.
  • Communication skills: You explain nuanced technical judgement clearly and help engineers apply it to their own products.

Requirements

  • 3+ years of relevant experience in LLM evaluation, ML evaluation, applied AI, data science, or a closely related field.
  • Direct experience designing and running evaluations for LLM-based applications or agents.
  • Experience designing evaluation suites, building datasets, and using results to influence or change a shipping decision.
  • Practical experience with LLM-as-a-judge approaches, including calibration against human judgement and understanding their limitations.
  • Strong experience in dataset construction and labelling, including sampling strategies, annotation guidelines, annotator agreement, contamination, and drift.
  • Strong grasp of measurement concepts, including sample size, statistical significance, variance between runs, and distinguishing real improvement from noise.
  • Working knowledge of LLM application architecture, including retrieval, tool calling, agent loops, prompts, and prompt management.
  • Experience advising or collaborating with engineering and product teams on evaluation design.
  • Strong written and spoken English.

What do we offer?

We focus on health, wealth, and empowering relationships:

  • Fully remote work with flexible work hours
  • Competitive salary
  • Opportunity to become a co-owner of the cooperative
  • Individual career development plan
  • Friendly team and company culture
  • Prioritization of mental and physical health in the workplace, with the freedom to make decisions about oneself, supported by peers committed to a healthy lifestyle.
  • Empowering relationships for engineering alongside colleagues who cherish growth mindsets in a unique environment that blends service and product craftsmanship.

What does the interview process look like?

  1. Initial Interview: We’ll start with a friendly 45-minute cultural and technical interview. Two members of our team will assess your cultural fit, past experience, engineering expertise, the major challenges you’ve tackled, and discuss your ideal workspace.
  2. You can choose between two Technical Deep Dive options:
    1. Homework Assignment: If there’s a match, we’ll provide a brief homework assignment designed to take around 2 hours to complete. This will be followed by a 1-hour technical interview to discuss the homework and conduct a technical deep dive.
    2. Pair Programming: If you prefer not to do a homework assignment, we’ll have a 2-hour technical deep dive session focused on a practical evaluation-design scenario.

Regardless of the outcome, we will provide you with constructive feedback to help you grow.

PEA Registration Number and Date of the Certificate 4003 / 10.10.2025

Stop Drowning in AI Hype

Get weekly insights from 50+ practitioners implementing AI in real businesses

Why You’ll Love It: