top of page

From Google Glass to Sign Language AI: Communication Choice Must Come First



Why captions, Sign Language-to-Text, wearable technology and human support should work together, not replace one another.



Featured-image alt text: A diverse group using sign language, caption glasses, mobility aids and human communication support in a public transport hub and café.
Featured-image alt text: A diverse group using sign language, caption glasses, mobility aids and human communication support in a public transport hub and café.

Featured-image alt text: A diverse group using sign language, caption glasses, mobility aids and human communication support in a public transport hub and café.


This is not only about me. It is about people, what they need and their right to communicate.


For a long time, communication technology has focused mainly on the human voice. We have speech recognition, audio description, translation and captions. Google Chrome can now create captions from audio playing in the browser.


These tools are useful, but Deaf people cannot simply do the same as hearing people. Some of us cannot access speech clearly without hearing aids, or even when we are wearing them. We are all different. One solution will never work for everybody.


I wore Google Glass around 2016

I believe it was around 2016 when I tried Google Glass connected by Bluetooth to a remote captioning service. Captions appeared privately in front of my eye.


At the time, I thought it was an interesting way to receive communication without having to keep looking down at another screen. The idea of caption glasses is not brand new, but the technology has moved forward a lot since then.


A few days ago, I watched a demonstration of Captify captioning glasses. They can use Wi-Fi, Bluetooth or a wired connection. I also know about SignGlasses. The captions are shown privately to the person wearing them, so my privacy concerns are different from my concerns about camera-based smart glasses.


Meta glasses are more complicated. A blind person might benefit from the camera, audio information or remote visual assistance. That could offer real independence. However, we still need to think seriously about data protection, facial identification, privacy and safeguarding children.


I am not simply for or against smart glasses. Access matters, but so do safety, privacy and personal choice.


What happened before 2020?

There has been much more public discussion about sign language since around 2020. Strictly Come Dancing helped bring BSL to a wider audience, and I now see sign language discussed regularly on Facebook and LinkedIn.


That is positive, but what happened before 2020?


Deaf people had already spent generations campaigning, teaching, interpreting, translating and fighting for recognition. Sign language did not suddenly appear when mainstream society noticed it.


The 1880 Milan Conference promoted oralism and contributed to sign language being suppressed in Deaf education. In the UK, the 1893 Elementary Education (Deaf and Blind Children) Act reinforced oralist teaching. Deaf children were expected to speak and lip-read, even when this did not meet their language needs.


The 1978 Warnock Report helped improve recognition of special educational needs. However, Deaf children were still often viewed mainly through disability and education, not through BSL, Deaf identity, language rights or the danger of language deprivation.


Later, we had the Disability Discrimination Act, Human Rights Act and Equality Act 2010. These laws mattered, but a law on paper does not automatically remove communication barriers.


BSL was not formally recognised by the UK Government as a language in its own right until 2003. The Milan resolutions were formally rejected in 2010, 130 years after they were adopted. The BSL Act finally became law in 2022.

This history is one reason why Deaf people ask difficult questions when a company introduces new sign-language technology. We know what can happen when decisions about our communication are made without us.


Why I follow this technology

For the past two years, I have independently followed Kara Technologies, Migam.ai, Signvrse, SilentSpeak, SignCaption, Sign-Speak, SignGlasses and TalkSign.


I am not employed by, paid by or representing these companies. I follow their work because I care about the future of communication. I want to know whether the technology genuinely helps people, who controls it and who benefits from it.


Over the past two weeks, sign-language users have not stopped discussing Google Pixel 11 and its Sign Language-to-Text technology. Hearing people have praised it too.


Well done to Google for bringing sign-language technology into a mainstream consumer product and a much wider conversation.


To clarify one point, Google did not sell Google Glass to Samsung. Google and Samsung are now working together on Android XR and new intelligent eyewear, alongside brands including Warby Parker and Gentle Monster.


What interests me about the movement

I am especially interested in MediaPipe Holistic because it can map landmarks across the face, hands and body.


Sign language does not come from the hands alone. Facial expression, mouth patterns, body position, movement and signing space can all affect the meaning.


A friend and I were signing about this. We compared it with creating a football video game such as FIFA.


Think about how many staff are involved in designing the stadiums and players, capturing movement and making a player respond naturally to a joystick. Even a small joystick movement must become a realistic movement on the screen.


Sign-language technology has an even more complicated job. A camera must follow the face, hands and body, understand how the movements work together and translate meaning from a three-dimensional signing space.


Google says its research used about 100,000 hours of sign-language video.


That scale is incredible. The camera is not just recording a person. The system is learning patterns across a huge landscape of sign language.


I find that exciting, but it also makes me ask questions.


SL2T starts with the person signing


Google DeepMind’s Sign Language-to-Text model currently allows a person to sign in ASL and produce English text on Pixel 11 through Gboard and Live Transcribe. Google says more devices and sign languages will follow. BSL is not supported yet.


This feels like an important change of direction. A lot of existing technology takes written or spoken information and generates sign language for a Deaf person to receive. SL2T starts with the Deaf person signing and helps them communicate outwards.


In the future, I would like to see both directions working together:

  • Sign language to text or speech

  • Speech or text to sign language


That could support a more complete two-way conversation. It is not about deciding that one company or system is best. Different tools can do different jobs.


AI must not replace interpreters and translators


I want to be very clear: AI must not replace qualified sign-language interpreters or translators.


However, it might help fill short, everyday gaps where arranging an interpreter is not practical.

This could include ordering food, asking something in a shop, visiting a front desk, discussing a car MOT, attending a sports match or asking for information at a transport hub.


Captions already fill some of these gaps.

SL2T could offer another choice.


I could sign naturally and have my message turned into text. Captions or speech-to-text could show me the other person’s reply.


That is very different from a court hearing, prison visit, hospital appointment, hospice discussion or safeguarding situation. In critical and complicated settings, we need qualified people who understand context, emotion, risk and responsibility.


Technology must never become an excuse for an organisation to refuse an interpreter, translator or another reasonable adjustment.


VRS is important, but direct communication matters too

Video Relay Services are important for telephone calls. When I use VRS, though, the interpreter does not automatically know where I am, why I am calling or what has happened. I might need to explain the context, main topic and purpose before the real conversation begins.


For a short and low-risk situation, SL2T might let me communicate directly without first explaining everything to an interpreter. That could give me more independence while allowing interpreters to focus on situations where their professional skills are essential.


Direct Video Calling is another useful choice. It allows a ASL/BSL user to contact an organisation directly by video instead of making a conventional voice call through a relay interpreter.


This is not a competition between AI and people. We need direct sign-language communication, interpreters and translators, VRS, Direct Video Calling, captions, SL2T, wearable technology and other human support.


The right option will depend on the person, the purpose, the risk and the situation.


My challenge to the technology industry


I encourage Apple, Samsung, Microsoft, Meta and others to build on Google’s progress. I do not mean copying a feature and rushing it to market. I mean investing responsibly in sign-language recognition, captions, wearable access and real communication choice.


Microsoft already has experience in accessibility, Teams captions, translation, AI, gaming and motion tracking. That knowledge could help sign-language communication, but Deaf people and sign-language experts must be involved from the start.


Competition can move technology forward, but accessibility must not become a race to be first.


I will continue asking:

Where did the training material come from?

Did the signers give informed permission?

Were they paid fairly?

Can they withdraw their data?

How will their identities and biometric information be protected?

Who tests the accuracy independently?

How can somebody report a harmful result?


Avatars might reduce some risks of identifying a real person, but they create other questions about accuracy, natural expression and accountability.


Interpreters and translators need protection too. Their recordings, appearance, signing and professional knowledge should not be used to train AI without clear permission, fair payment and agreed limits.


My focus is the future

These are my personal thoughts. I do not have every answer. I am still learning, watching developments and asking questions.


Technology should remove communication barriers, not build another divide between Deaf and hearing people. This is not about one group controlling how another group communicates. It is about equality, choice and people understanding one another.


New technology must be developed with Deaf people, not simply shown to us after all the important decisions have been made.


My focus is the future: young people being able to communicate through a real choice of technology, captions, sign language, interpreters, translators and other human support.


The goal is not to make everybody communicate in the same way.


The goal is to give every person an equal, safe and meaningful way to communicate.


Further reading

Disclosure: I am not employed by, paid by or representing the technology companies named in this article. These are my own views.

 
 
 

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page