Back to Jobs
Pencil
AI & Machine Learning Just now

Quality & Evals Lead (Generative AI Creative Measurement)

Pencil
United StatesUnited States
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Prompt Engineering17mFree Trial ✨
Start 10-Day Free Trial
Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
Python ScriptingNLPChatbot Dev

Want to know if you're a match for this job?

Calculate My Match Score

At Pencil, we are driving innovation in advertising technology through our state-of-the-art SaaS product, which harnesses Generative AI to redefine content creation. Our mission is to make AI the default in advertising without replacing creative people.

We aim to ensure our technology isn’t just for big brands but also helps small businesses and creative individuals.

The Role:

Pencil’s agents produce advertising creative at scale—video, adaptations, and multiple formats—orchestrated by our agentic system, Scribble. The product’s success depends on output quality being right the first time. This role focuses on measuring and improving quality.

You’ll oversee evaluations end-to-end. Internally, this involves systems measuring whether agents, skills, and workflows produce client-ready work. Externally, it includes demonstrating quality to customers and the market through evidence behind client RFP responses, creative enablement packages, and adevals.ai, our public evaluation site.

Measuring creative quality is challenging due to subjective standards. Your task is to build reliable measurement systems: rubrics defining what “good” means, calibration between human reviewers and automated judges, and a golden dataset that remains representative as the product and client base evolve. You’ll then integrate these metrics into the business’s operations and sales strategy.Key Responsibilities:

  • Define the charter for Evals & Quality, including the problem, success metrics, and scope.
  • Establish core quality metrics like first-pass-right, coverage, and regression detection, driving their improvement.
  • Develop judging methodologies, including rubrics for creative quality and calibration processes to align automated judges and human raters over time.
  • Design and implement human QC mechanisms, ensuring human reviews feed into automated systems and scale effectively.
  • Curate and maintain the golden dataset, ensuring coverage across formats and client contexts, with proper versioning.
  • Build eval loops for other teams to use independently, enabling self-service evaluation setups.
  • Own adevals.ai as a product, ensuring high design quality and credibility.
  • Provide evidence for commercial work, including client RFP responses and creative enablement packages.
  • Collaborate with Agent Architects to set client-specific quality bars via configuration and rubrics.
  • Manage the full product lifecycle: discovery, launch, and post-launch audits, with written hypotheses before each ship and reviews 14 days post-launch.

Your Background:

Ideal candidates can articulate:

  • A quality, evals, or ML measurement problem they owned as a product, complete with roadmap and users.
  • Measurement systems they shipped in subjective domains, detailing how they defined quality and maintained reliability at scale.
  • Direct experience with golden datasets, rubric design, judge calibration, and inter-rater reliability, including past failures and resolutions.
  • A product decision derived from evidenced customer needs in a domain they had to learn.
  • A metric they owned that influenced decisions outside their team.
  • Experience presenting or defending methodology to clients.

You’ll Thrive Here If You...

  • Move Fast: Comfortable making decisions with imperfect information, iterating quickly, and resolving issues before they escalate.
  • Stay Human: Prioritize empathy for users and teammates, designing with real workflows in mind.
  • Act Like an Owner: Take responsibility for outcomes, not just outputs. Proactively identify and fix problems.
  • Keep It Simple: Eliminate unnecessary complexity, design reusable patterns, and make hard things intuitive.
  • Love to Surprise: Delight users with unexpected improvements, especially in polish and usability.

KPIs & Success Measures:

  • First-pass-right: Share of generations a client would ship without rework (lead metric).
  • Eval coverage: Share of skills and formats with live eval loops.
  • Judge-human calibration: Agreement between automated judges and human raters, maintained over time.
  • Golden dataset health: Representative, versioned, and current across formats and client contexts.
  • Commercial evidence: adevals.ai live; evals used in RFP wins and enablement packages.
  • Learning velocity: Every launch reviewed 14 days post-launch against its stated hypothesis.

How would you rate this job post?

See what other professionals think about this role.

banner

Pencil (operating at trypencil.com) is an enterprise-grade AI marketing platform engineered for automation and governance in creative and media workflows. Founded by James Chadwick and Sumukh Avadhani and headquartered in London, Pencil redefines marketing operations by aggregating AI models into a unified operating system. Unlike traditional tools that fragment workflows or lack scalability, Pencil enforces enterprise-grade governance while turning production efficiencies into measurable media growth. Under the hood, the platform consolidates disparate AI agents, ensuring compliance, version control, and seamless collaboration across briefing, creative generation, and performance tracking. This empowers global brands—including Diageo, Unilever, and L’Oréal—to accelerate campaign execution, reduce manual overhead, and maintain consistency at scale. Backed by $4.1M in funding across two rounds, including a Seed round in 2020, Pencil has also been recognized as a Gartner Cool Vendor for its innovative approach to agentic marketing automation.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More