Google DeepMind released SL2T today, a sign-language-to-text translation model that ships inside Gboard and Live Transcribe on the Pixel 11, marking the first time sign language AI has landed inside a mainstream consumer product. The model translates American Sign Language into English text in a streaming, on-device pipeline, letting Deaf users sign to their phone anywhere they would otherwise type. It scores 70 BLEURT zero-shot on the FLEURS-ASL benchmark, a number Google DeepMind says is significantly higher than any previously reported result.
SL2T was trained on more than 100,000 hours of data across over 50 sign languages, with roughly a quarter of the corpus in ASL. Google DeepMind reports that jointly training across diverse languages, dialects, and signer proficiency levels caused the model to outperform single-language variants in internal experiments. ASL to English is the launch pair; additional languages and devices are on the roadmap.
The consumer surfaces are the point. In Gboard, users sign to search the web, draft messages, or hand queries to Gemini. In Live Transcribe, they sign responses inside conversations instead of typing back and forth. The company frames this as parity with the dictation experience hearing users have taken for granted for years.
“According to our testers, signing in ASL is faster, more natural, and more delightful than typing in English.”— Google DeepMind Sign Language Team, Authors of the SL2T announcement
Key facts
- 01SL2T is trained on 100,000+ hours of data across more than 50 sign languages, with about a quarter of that data in American Sign Language.
- 02The model scores 70 BLEURT zero-shot on the FLEURS-ASL (sd-test) benchmark, which Google DeepMind says is significantly higher than any previously reported score.
- 03SL2T powers sign-to-text dictation in Gboard and Live Transcribe on Pixel 11, launching with ASL to English first.
- 04There are 70 million Deaf and hard of hearing people worldwide who use one of more than 200 sign languages.
- 05The system tracks pose landmarks on-device via MediaPipe Holistic and discards raw video, addressing privacy concerns.
Sign languages are not transliterated speech. They carry their own grammar, lexicon, and non-manual markers conveyed through simultaneous movements of the hands, arms, torso, head, and face. That is why earlier attempts like sign language gloves failed on contact with real users: they treated signing as a hand-motion problem rather than a full language-translation problem.
SL2T is designed for both halves. An on-device MediaPipe Holistic model tracks pose landmark locations on the signer's body, and only those geometric coordinates are sent to the server. The raw camera feed is discarded immediately, a privacy design decision that also cuts bandwidth. The server model then translates the coordinate sequence directly into English text, skipping the intermediate gloss annotations that limited prior academic work.
That direct-from-landmarks approach removes artificial vocabulary caps and lets translation quality scale with data rather than with a hand-tuned symbol table. On FLEURS-ASL sd-test, the 70 BLEURT zero-shot score reflects fluent handling of complex constructions, though errors persist on rare signs, rapid fingerspelling, passive voice, and tense without surrounding context. In one sample the model rendered ASL for a Cook Islands travel passage as "The Cook Islands have no cities and consist of 15 islands. The two main islands are Rarotonga and Aitutaki" — close to the original English source.
Beyond benchmark scores, Google DeepMind says it invested in the failure modes that matter in production: streaming latency, hallucination suppression when no one is signing, performance for the 10% of signers who are left-handed, and one-handed signing accuracy for users holding the phone in the other hand. Those are the details that separate a demo from a shipped feature.
The project was conceived by Sam Sepah, a Deaf Googler, and developed alongside the AI Sign Language Advisory Committee, a group Google DeepMind assembled from Deaf organizations and subject-matter experts to guide development priorities. The company published a joint impact report with the committee for the SL2T 1.0 launch, detailing capabilities and current limitations, and says it will repeat that process for future releases.
Limitations are real. SL2T covers ASL to English at launch, which leaves the vast majority of the roughly 70 million sign language users worldwide waiting. The model still stumbles on rare signs and rapid fingerspelling, and Google DeepMind has not disclosed word-error rates by signer demographic beyond noting the left-handed and one-handed work. Deployment is initially tied to a single flagship phone, and expansion to other Android devices is described only as "coming soon."
SL2T is one of the clearest examples in recent memory of frontier AI capability being pointed at a use case where the baseline experience — typing English as a second language — is genuinely worse than what the model can now provide. The economics for Google are secondary here: the Pixel 11 gets a differentiated accessibility story, Gemini gets a new input modality, and Android gets a template for rolling sign input to more languages and more devices. If SL2T holds up outside curated demos, it also raises the floor for what accessibility means on a smartphone in 2026, and puts pressure on Apple and Samsung to answer with something comparable rather than a captioning feature.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




