Back to Careers
AI Evaluation Engineer
Build evaluation harnesses, test datasets, and release gates that prove production AI systems behave correctly.
Remote / HybridSecurity, Quality & ReliabilityFull-time
Role Overview
The AI Evaluation Engineer is responsible for measuring whether Tactical Edge's agents, workflows, retrieval systems, and customer solutions are accurate, safe, and reliable before they reach production. You will turn expected behavior into eval suites, build regression datasets, test failure modes, and help teams decide when an AI system is ready to ship.
What You'll Do
What We're Looking For
How We Work
Outcome-driven
Measured AI quality over vibes
Enterprise-first
Quality, trust, and predictable behavior
Evaluation by design
Release checks built before launch
Small teams, high ownership
Autonomy with accountability
What You'll Get
- Ownership of quality for production AI systems
- Exposure to enterprise-scale platforms and deployments
- Cross-functional collaboration with product, AI, and engineering teams
- Competitive compensation (role/location dependent)
- Flexible work setup where applicable
Hiring Process
- 1Intro call (context + fit)
- 2AI evaluation scenario discussion (eval strategy for a real system)
- 3Cross-functional interview (engineering/product perspective)
- 4Final conversation
We value measured judgment, clear eval design, and practical release discipline.