The machines that learned to listen
Science Photo LibraryVoice recognition technology makes many aspects of modern life easier. The seeds were sown a lot further back than you might think.
A toddler meanders unsteadily through the living room, pausing by a sleek black cylinder in the corner. âAlexa,â he says in a high-pitched voice. âPlay children music.â The cylinder acknowledges the request, despite the muffled pronunciation, and the music starts.
Alexa, a cloud-based speech recognition software from Amazon and the brain of its black cylindrical loudspeaker Echo, has been a big hit around the world â except for the younger ones, who take it for granted. Children will grow up alongside it, just as Alexa will evolve, as the AI powering it learns to answer more and more questions, and â Â perhaps â one day even converses freely with people.
But anyone older than 10 will know that it hasnât always been like that. Speech recognition software has come a long, long way to where we are today. Echo is slimmer than a beer glass, but the first speech recognition machines â developed during the middle of the 20th Century â nearly took up an entire room.
AmazonHumans have long wanted to speak to machines â or at least make them talk to us. âVoice enables unbelievably simple interaction with technology â the most natural and convenient user interface, and the one we all use every day,â says Jorrit Van der Meulen, VP at Amazon Devices and Alexa EU. âVoice is the future.â
Back in 1773, Russian scientist Christian Kratzenstein, a professor of physiology in Copenhagen, seemed to be thinking along the same lines. He built a peculiar device that produced sounds similar to human vowels using resonance tubes connected to organ pipes. Just over a decade later, Wolfgang von Kempelen in Vienna created a similar Acoustic-Mechanical Speech Machine. And in the early 19th Century, English inventor Charles Wheatstone improved on von Kempelen's system with resonators made out of leather. Their configuration could be changed or controlled by hand to produce different speech-like sounds.
Then in 1881, Alexander Graham Bell, his cousin Chichester Bell and Charles Sumner Tainter built a rotating cylinder with a wax coating, with a stylus that would cut vertical grooves, responding to incoming sound pressure. The invention paved the way for the first recording machine, the "Dictaphone", patented in 1907. The idea was to get rid of stenographers by using the machine to record dictation of notes and letters for a secretary, so that they could later be typed offline. The invention took off, with more and more offices around the globe sporting a secretary with a clunky earpiece, listening to the recordings and transcribing them.
But all those baby steps kept machines passive â until âAudreyâ, the Automatic Digit Recognition machine, came along in 1952. Made by Bell Labs, the huge machine occupied a six-foot-high relay rack, consumed substantial power and had streams of cables. It could recognise the fundamental units of speech sounds, which are called phonemes.
Back then, computing systems were extremely expensive and inflexible, with limited memory and computational speed. But regardless, Audrey could recognise the sound of a spoken digit â zero to nine â with more than 90% accuracy, at least when uttered by its developer HK Davis. It worked with 70-80% accuracy for a few other designated speakers, but far less well with voices it was unfamiliar with. âThis was an amazing achievement for the time, but the system required a room full of electronics, with specialised circuitry to recognise each digit,â says Charlie Bahr of Bell Labs Information Analytics.
Science Photo LibraryBecause Audrey could recognise only voices of designated speakers, its use was limited: for instance, it could offer voice dialling by, say, toll operators, but it wasnât really a necessity because in most cases manual push-button dialling of numbers was cheaper and easier. Audrey was an early bird â it preceded general purpose computers, and although it was not used in production systems, âit showed that speech recognition could be made practicalâ, says Bahr.
But there was another goal. âI believe Audrey was initially developed to reduce bandwidth, the volume of data travelling over the wires,â says Bahrâs colleague Larry OâGorman of Nokia Bell Labs. Recognised speech would require much less bandwidth than the original sound waves. But as telephone switches became digital in the 1970s and 80s, they enabled faster and cheaper call routing, while staying dependent upon an operator recognising a personâs request to dial a number. So, in the 1970s and 80s, a huge effort in Bell Labsâ speech research was to simply do the following: recognise zero to nine digits, and âyesâ or ânoâ. âWith recognition of these 12 words, the telephone system was able to complete the transition to machine-only telephony,â says OâGorman.
Audrey was not the only kid on the block, though. In the 1960s, several Japanese teams worked on speech recognition, with the most notable ones a vowel recogniser from the Radio Research Lab in Tokyo, a phoneme recogniser from Kyoto University, and a spoken-digit recogniser from NEC Laboratories.
At the 1962 World Fair, IBM showcased its "Shoebox" machine, able to understand 16 spoken English words. There were other efforts in the US, UK and the Soviet Union, with Soviet researchers inventing the dynamic time-warping (DTW) algorithm that they used to build a recogniser capable of working with a 200-word vocabulary. But all these systems were mostly based on template matching, where individual words are matched against stored voice patterns.
The most significant leap forward of the time came in 1971, when the US Department of Defenseâs research agency Darpa funded five years of a Speech Understanding Research programme, aiming to reach a minimum vocabulary of 1,000 words. A number of companies and academia including IBM, Carnegie Mellon University (CMU) and Stanford Research Institute took part in the programme. Thatâs how Harpy, built at CMU, was born.
Unlike its predecessors, Harpy could recognise entire sentences. âWe donât want to look things up in dictionaries â so I wanted to build a machine to translate speech, so that when you speak in one language, it would convert what you say into text, then do machine translation to synthesise the text, all in one,â says Alexander Waibel, a computer science professor at Carnegie Mellon who worked on Harpy and another CMU machine, Hearsay-II.
iStockMoving from single words to phrases wasnât easy. âWith sentences, you get words flowing into each other, you get a lot of confusion and donât know where the words end and where they begin. So you have things like âeuthanasiaâ, which could be âyouth in Asiaâ,â says Waibel. âOr if you say âGive me a new displayâ it could be understood as âgive me a nudist playââ.â
All in all, Harpy recognised 1,011 words â approximately the vocabulary of an average three-year-old â with reasonable accuracy, thus  achieving Darpaâs original goal. It âbecame a true progenitor to more modern systemsâ, says Jaime Carbonell, director of the Language Technologies Institute at CMU, being âthe first system that successfully used a language model to determine which sequences of words made sense together, and thus reduce speech recognition errorsâ.
In the years that followed, speech recognition systems evolved further. In the mid 1980s, IBM built a voice activated typewriter dubbed Tangora, capable of handling a 20,000-word vocabulary. IBMâs approach was based on a hidden Markov model, which adds statistics to digital signal processing techniques. The method makes it possible to predict the most likely phonemes to follow a given phoneme.
IBMâs competitor Dragon Systems came up with its own approach, and technological advances finally pushed speech recognition far enough that it could find its first applications â such as dolls that kids could train to speak. But still, despite these successes, all the programs at the time used discrete dictation, meaning the user had to pause⊠after⊠every⊠word. In 1990, Dragon released the first consumer speech recognition product, Dragon Dictate, for a whopping $9,000. Then in 1997 Dragon NaturallySpeaking appeared â the first continuous speech recognition product.
âBefore that time, speech recognition products were limited to discrete speech, meaning that they could only recognise one word at a time,â says Peter Mahoney, senior vice president and general manager of Dragon, Nuance Communications. âBy pioneering continuous speech recognition, Dragon made it practical for the first time to use speech recognition for document creation.â Dragon NaturallySpeaking recognised speech at about 100 words per minute â and it is still used today, for instance, by many doctors in the US and the UK to document their medical records.
iStockIn the last 10 years or so, machine learning techniques loosely based on the workings of the human brain have allowed computers to be trained on huge datasets of speech, enabling excellent recognition across many people using many different accents. Â
Still, the technology stalled until Google released its Google Voice Search app for the iPhone. Googleâs trick was to use cloud computing to process the data received by its app. Suddenly, publicly available voice recognition had massive amounts of computing power at its disposal. Google was able to run large-scale data analysis for matches between the user's words and the huge number of human-speech examples it had amassed from billions of search queries. In 2010, Google added "personalised recognition" to Voice Search on Android phones, and Voice Search to its Chrome browser in mid-2011. Apple quickly offered its own version, called Siri, while Microsoft called its AI Cortana, named after a character in the popular Halo video game franchise.
So whatâs next? âWithin speech processing, the most mature technology is speech synthesis,â says OâGorman. âMachine voices now are largely indistinguishable from a humanâs. But automatic speech recognition is still far less successful than the human ear in many situations.â While speech can be automatically recognised by a clearly speaking person in an environment with little noise, the so-called âcocktail-party effectâ â where humans can understand a single speaker in the din of a party â is still beyond any state-of-the-art technology. Even with Alexa, in a noisy room you have to make sure youâre right near the black cylinder and speak to it clearly and loudly.
Amazonâs attempt at voice recognition was inspired by the Star Trek computer, says Van der Meulen, with the aim of creating a computer in the cloud thatâs controlled entirely by your voiceâso that you could converse with it in a natural way. Sure, the magic of Hollywood still has the edge on todayâs technology, but, says Van der Meulen, âweâre in a golden age of Machine Learning and AI. Weâre still a long way from being able to do things the way humans do things, but weâre solving unbelievably complex problems every day.â
If you liked this story, sign up for the weekly bbc.com features newsletter, called âIf You Only Read 6 Things This Weekâ. A handpicked selection of stories from BBC Future, Earth, Culture, Capital, Travel and Autos, delivered to your inbox every Friday.
Â
