Bounded voice decisions for consequential inputs.
Choose either peer capability: verify a consequential action, or compare a candidate against a preaccepted specification. AssemblyAI supplies what was heard; each capability makes its own bounded decision.
SPEECH ≠ VERBATIM TRANSCRIPT ≠ CLEAN DICTATION ≠ APPLICATION-VALID INPUT ≠ SPECIFICATION MATCH ≠ SPEAKER SIMILARITY ≠ VERIFIED INPUT ≠ AUTHORIZED ACTION
Bounded payment input
Dictate a payment instruction
Example: Pay Acme Labs fifteen thousand dollars against invoice INV-14892 from Growth next Friday.
Compare this instruction independently
Speak a payment instruction for this capability alone. Its Dictation result and specification comparison do not change capability 01.
payment-approval-001 v1
spec hash: b3f0d50ef7cc754ab11df7df007cda604ea115af4a422e50d8f2ce2f52747656
This is the accepted reference. Certain compares canonical typed values, not sentence strings.
Candidate versus accepted values
No candidate has been compared in this workspace yet. Use the candidate input above to produce an independent result.
A specification match does not satisfy amount repeat or invoice confirmation, and it does not authorize a downstream action.
Correct hearing, invalid input
Say, or simulate: “Pay Acme Labs $50,000 against invoice INV-14892 from Growth next Friday.” AssemblyAI may transcribe $50,000 perfectly. Certain must still answer BLOCKED because the contract caps amounts at $25,000. The specification must also answer MISMATCH. No confirmation can promote a blocked payload to VERIFIED.
FEEDBACK / FOUNDRY EVIDENCEWhat was tested outside the product flowBOUNDED
These findings are kept separate from the two capabilities. They describe AssemblyAI substrate behavior and evidence limits, not additional product features.
- Dictation semantics: request ordering, missing config, unsupported audio, cleanup output, and authentication behavior were exercised with retained response evidence.
- Human corpus: 23 supported English fixtures were run with fixed audio hashes. Results were 19/50 exact and 27/50 normalized target matches. This is corpus-bounded quality evidence, not a general accuracy claim.
- Context comparison: matched prompt and keyterm arms reused identical audio and retained separate scorecards. Observed improvements remain bounded quality/DX findings, not confirmed AssemblyAI bugs.
- Limits: speaker similarity is not shipped, and the newly added challenge plus exact $50k sentence were not captured with a physical microphone in the agent environment.
Full retained evidence: research/assemblyai/FINAL_FOUNDRY_REPORT.md