
Evaluation Framework
Build the framework that proves our AI system's answers are right, cited, and honest.

Why this exists
Genie answers questions in domains where a wrong answer carries real cost. Our flagship deployment covers workplace rights across 22,000+ US jurisdictions, for people who may be in a hard moment when they ask. Every answer must be right, cited, and honest about what the system doesn't cover. A system held to that bar needs a framework that proves, before every release, that it still clears it.
What you'll build
An evaluation framework for a production AI system.
Test sets that define what correct means.
Checks that an answer actually follows from the sources it cites.
Scoring for abstention. In our domains, a confident guess is worse than saying “not covered.”
Automated grading calibrated against human judgment.
Gates wired into our release process.
This brief is deliberately high level. The deeper context on our corpus, tooling, and failure modes is shared once you're selected.
What done looks like
Releases gated on your framework. A red result stops a ship, and the team trusts it enough to let it.
Who this fits
Engineers whose first question is “how do we know it works.” People who like building the thing that keeps everyone else honest.
The bounty
paid on completion.