Recognition
Can it recognize an unseen signer?
Random clip splits can leak signer-specific patterns into training. The evaluation needed to hold out complete signers and report results across every signer, not only the strongest fold.
Client A · Sign language AI · Case study
Client A had consented sign-language data, domain expertise and a working research direction. The next question was harder: how do we prove that recognition works for unseen signers and that generated motion preserves the sign?
The problem
Recognition and generation were being discussed together, even though they fail differently and require different evidence. The case study separated them before assigning any readiness score.
Recognition
Random clip splits can leak signer-specific patterns into training. The evaluation needed to hold out complete signers and report results across every signer, not only the strongest fold.
Generation
Visual polish can hide changed hand paths, timing, orientation or non-manual features. The output needed motion evidence and Deaf review, not a subjective “looks good” decision.
Generation pilot
Several rendering approaches were tested against the same short motion case. The selected candidate produced the strongest observed combination of motion retention, articulated hands, a stable face and a visible mouth.
Selected pilot output for evaluation case G_000633. Generated media shown with permission.
Benchmark definition
The benchmark is designed to prevent a strong aggregate number or polished video from hiding failures that matter to actual sign-language use.
AgentOps assessment
The review connects each readiness claim to evidence, identifies the current gap and keeps a successful artifact from becoming an unsupported production claim.
Research, recognition, generation and approval boundaries are separated.
In progressConsented, pseudonymized data and ownership boundaries are documented.
In progressSonZo components, external renderers and validation responsibilities are traceable.
PartialUncertain recognition and semantically altered output require explicit handling.
In progressThe LOSO protocol is defined; full seven-fold execution and generation smoke testing remain.
In progressRelease thresholds, rollback rules and final approvers still need to be frozen.
Not readyInputs, prompts, hashes, diagnostics and outputs exist; one run registry is still needed.
PartialCurrent decision
The recognition protocol is defined but not yet fully executed. The selected generation output is a reproducible pilot candidate, not proof that the broader vocabulary will preserve meaning.
The next evidence comes from all seven signer-disjoint recognition folds, a fixed generation smoke set, measurable acceptance thresholds and a reviewer decision trail.
The Agent Production Readiness Assessment shows what your system can prove today, what remains blocked and who must own the next decision.
Book a Production Readiness Review