Client A · Sign language AI · Case study

Turning a promising AI demo into measurable evidence.

Client A had consented sign-language data, domain expertise and a working research direction. The next question was harder: how do we prove that recognition works for unseen signers and that generated motion preserves the sign?

7 pseudonymized signers7-fold signer-disjoint evaluation7 production gates

The problem

A good-looking output is not the same as a validated system.

Recognition and generation were being discussed together, even though they fail differently and require different evidence. The case study separated them before assigning any readiness score.

Recognition

Can it recognize an unseen signer?

Random clip splits can leak signer-specific patterns into training. The evaluation needed to hold out complete signers and report results across every signer, not only the strongest fold.

Generation

Did the rendered avatar preserve the sign?

Visual polish can hide changed hand paths, timing, orientation or non-manual features. The output needed motion evidence and Deaf review, not a subjective “looks good” decision.

Generation pilot

The first reproducible candidate

Several rendering approaches were tested against the same short motion case. The selected candidate produced the strongest observed combination of motion retention, articulated hands, a stable face and a visible mouth.

Selected pilot output for evaluation case G_000633. Generated media shown with permission.

  • InputVersioned 4-second motion-control video
  • RendererSeedance 2.5 through the Runway API
  • ReproducibilityModel, seed, prompt, task record and file hashes retained
  • Current meaningSuccessful pilot candidate, not a completed benchmark result

Benchmark definition

Two evaluations with explicit boundaries

The benchmark is designed to prevent a strong aggregate number or polished video from hiding failures that matter to actual sign-language use.

Macro-F1Primary SLR metric across eligible signs
7-fold LOSOEach signer becomes the untouched test signer once
CalibrationConfidence and abstention are evaluated, not assumed
Motion fidelityNormalized joint and hand trajectories compared with control
Non-manualsMouth, eyebrows, gaze and head behavior measured separately
Deaf reviewSemantic preservation and intelligibility remain human decisions

AgentOps assessment

The Seven Production Gates review

The review connects each readiness claim to evidence, identifies the current gap and keeps a successful artifact from becoming an unsupported production claim.

01
Purpose and Scope

Research, recognition, generation and approval boundaries are separated.

In progress
02
Data and Evidence

Consented, pseudonymized data and ownership boundaries are documented.

In progress
03
Architecture and Tool Authority

SonZo components, external renderers and validation responsibilities are traceable.

Partial
04
Safety and Governance

Uncertain recognition and semantically altered output require explicit handling.

In progress
05
Evaluation and Validation

The LOSO protocol is defined; full seven-fold execution and generation smoke testing remain.

In progress
06
Deployment and Release

Release thresholds, rollback rules and final approvers still need to be frozen.

Not ready
07
Observability and Continuous Improvement

Inputs, prompts, hashes, diagnostics and outputs exist; one run registry is still needed.

Partial

Current decision

Evidence is advancing. Production approval is not implied.

The recognition protocol is defined but not yet fully executed. The selected generation output is a reproducible pilot candidate, not proof that the broader vocabulary will preserve meaning.

The next evidence comes from all seven signer-disjoint recognition folds, a fixed generation smoke set, measurable acceptance thresholds and a reviewer decision trail.

Return to the 7 Production Gates

Have a working AI system but no defensible production decision?

The Agent Production Readiness Assessment shows what your system can prove today, what remains blocked and who must own the next decision.

Book a Production Readiness Review