Tactical Edge
Back to Careers

AI Evaluation Engineer

Build evaluation harnesses, test datasets, and release gates that prove production AI systems behave correctly.

Remote / HybridSecurity, Quality & ReliabilityFull-time

Role Overview

The AI Evaluation Engineer is responsible for measuring whether Tactical Edge's agents, workflows, retrieval systems, and customer solutions are accurate, safe, and reliable before they reach production. You will turn expected behavior into eval suites, build regression datasets, test failure modes, and help teams decide when an AI system is ready to ship.

What You'll Do

  • Design and maintain eval suites for agents, RAG workflows, tool calls, and customer-facing AI features.
  • Build golden datasets, adversarial test cases, and regression checks from real customer workflows.
  • Validate functional behavior, reasoning quality, retrieval quality, safety boundaries, and failure handling.
  • Collaborate with engineering, product, AI, and delivery teams early in the development lifecycle.
  • Define release gates that combine automated tests, human review, telemetry, and quality thresholds.
  • Analyze production issues and convert them into durable eval coverage.
  • Contribute to evaluation tooling, dashboards, and repeatable quality processes.
  • Document what the system is expected to do, how it is tested, and where known limits remain.
  • What We're Looking For

  • Experience testing production software, AI systems, data-heavy workflows, or distributed systems.
  • Strong understanding of test design, regression testing, and quality measurement.
  • Ability to turn ambiguous AI behavior into concrete pass/fail criteria and scored rubrics.
  • Comfort working with developers, product managers, AI engineers, and delivery teams.
  • Attention to detail combined with pragmatic judgment about what matters in production.
  • Familiarity with test automation, CI/CD pipelines, eval harnesses, or observability tools.
  • Bonus: Experience with LLM evals, RAG evaluation, prompt regression testing, or human-in-the-loop review.
  • How We Work

    Outcome-driven

    Measured AI quality over vibes

    Enterprise-first

    Quality, trust, and predictable behavior

    Evaluation by design

    Release checks built before launch

    Small teams, high ownership

    Autonomy with accountability

    What You'll Get

    • Ownership of quality for production AI systems
    • Exposure to enterprise-scale platforms and deployments
    • Cross-functional collaboration with product, AI, and engineering teams
    • Competitive compensation (role/location dependent)
    • Flexible work setup where applicable

    Hiring Process

    1. 1Intro call (context + fit)
    2. 2AI evaluation scenario discussion (eval strategy for a real system)
    3. 3Cross-functional interview (engineering/product perspective)
    4. 4Final conversation

    We value measured judgment, clear eval design, and practical release discipline.