neural networks -- 7/17/24
Today's selection-- from Genius Makers by Cade Metz. Work on neural networks, one of the early steps in the rapidly advancing discipline of artificial intelligence:
“In 1995, two Bell Labs researchers—Vladimir Vapnik and Larry Jackel—made a bet. Vapnik said that within ten years, ‘no one in their right mind would use neural nets.’ Jackel sided with the connectionists. They bet a ‘fancy dinner,’ with [researcher Yann] LeCun serving as witness when they typedup the agreement and signed their names. Almost immediately, it started to look as though Jackel would lose. As the months passed, another chill settled over the wider world of connectionist research. Pomerleau's truck could drive itself. Sejnowski's NETtalk could learn to read aloud. And LeCun's bank scanner could read handwritten checks. But it was clear the truck couldn't deal with anything more than private roads and the straight lines of a highway. NETtalk could be dismissed as a party trick. And there were other ways of reading checks. LeCun's convolutional neural networks didn't work when analyzing more complex images, like photos of dogs, cats, and cars. It wasn't clear if they ever would. As it turned out, Jackel eventually won the bet, but it was a hollow victory. Ten years after their wager, researchers might still have been using neural nets, but the technology couldn't do that much more than it had done on LeCun's desktop machine all those years before. ‘I won the bet mostly because Yann didn't give up,’ Jackel says. ‘He was largely ignored by the outside community, but he didn't give up.’
“Not long after the bet was settled, during a lecture on artificial intelligence, a Stanford University computer science professor named Andrew Ng described neural networks to a roomful of graduate students. Then he added a caveat. ‘Yann LeCun,’ he said, ‘is the only one who can actually get them to work.’ But even LeCun was unsure of the future. With a few wistful words on his personal website, he wrote about his chip research as something stuck in the past. He described the silicon processors he helped build in New Jersey as ‘the first (and maybe the last) neural net chips to actually do something useful.’ Years later, when asked about these words, he dismissed them, quickly pointing out that he and his students had returned to the idea by the end of the decade. But the uncertainty he felt was there on the page. Neural networks did need more computing power, but no one realized just how much they needed. As Geoff Hinton later put it: ‘No one ever thought to ask: “Suppose we need a million times more?”’
“While Yann LeCun was building his bank scanner in New Jersey, Chris Brockett was teaching Japanese in the Department of Asian Languages and Literature at the University of Washington. Then Microsoft hired him as an AI researcher. The year was 1996, not long after the tech giant created its first dedicated research lab.
“Microsoft aimed to build systems that could understand natural language-the everyday way people write and talk. At the time, this was the work of linguists. Language experts like Brockett, who had studied linguistics and literature in his native New Zealand and later in Japan and the U.S., spent their days writing detailed rules meant to show machines how humans pieced their words together. They would explain why ‘time flies,’ carefully separate ‘contract’ the noun from ‘contract’ the verb, describe in minute detail the strange and largely unconscious way that English speakers choose the order of their adjectives, and so on. It was a task reminiscent of the old Cyc project in Austin or the driverless car work at Carnegie Mellon before Dean Pomerleau came along—an effort to re-create human knowledge that wouldn't reach its endgame for decades, no matter how many linguists Microsoft hired. In the late '90s, following the lead of prominent researchers like Marvin Minsky and John McCarthy, this is how most universities and tech companies built computer vision and speech recognition as well as natural language understanding. Experts pieced the technology together one rule at a time.
![]() |
| An artificial neural network |
“Sitting in an office at Microsoft headquarters just outside Seattle, Brockett spent nearly seven years writing the rules of natural language. Then, one afternoon in 2003, inside an airy conference room down the hall, two of his colleagues unveiled a new project. They were building a system that translated between languages using a technique based on statistics—how often each word appeared in each language. If a set of words appeared with the same frequency and the same context in both languages, that was the likely translation. The two researchers had started the project only six weeks earlier, and it was already producing results that looked at least a little like real language. As he watched the presentation, sitting at the back of the crowded room, perched atop a long row of trash cans, Brockett had a panic attack—which he thought was a heart attack—and was rushed to the hospital. He later called it his ‘come-to-Jesus moment,’ when he realized he had spent six years writing rules that were now obsolete. ‘My fifty-two-year-old body had one of those moments when I saw a future where I wasn't involved,’ he says.
“The world's natural language researchers soon overhauled their approach, embracing the kind of statistical models unveiled that afternoon at the lab outside Seattle. This was just one of many mathematical methods that spread across the larger community of AI researchers in the 1990s and on into the 2000s, with names like ‘random forests,’ ‘boosted trees,’ and ‘support vector machines.’ Researchers applied some to natural language understanding, others to speech recognition and image recognition. As the progress of neural networks stagnated, many of these other methods matured and improved and came to dominate their particular corners of the AI landscape. They were all a (very) long way from perfection. Though the early success of statistical translation had been enough to send Chris Brockett to the hospital, it worked only to a point, and only when applied to short phrases—pieces of a sentence. Once a phrase was translated, a complex set of rules was needed to get it into the right tense and apply the right word endings and line it up with all the other short phrases in a sentence. Even then, the translation was jumbled and only vaguely correct, like that childhood game where you build a story by rearranging little slips of paper holding just a handful of words. But this was still beyond what a neural network could do. By 2004, a neural network was seen as the third best way to tackle any task—an old technology whose best days were behind it. As one researcher told Alex Graves, then a young graduate student studying neural networks in Switzerland: ‘Neural networks are for people who don't understand stats.’ While hunting for a major at Stanford, a nineteen-year-old undergraduate named Ian Goodfellow took a class in what was called cognitive science—the study of thought and learning—and at one point, the lecturer dismissed neural networks as a technology that couldn't handle ‘exclusive-or.’ It was a forty-year-old criticism debunked twenty years earlier.
“In the United States, connectionist research nearly vanished from the top universities. The one serious lab was at New York University, where Yann LeCun took a professorship in 2003, his hair pulled back in a ponytail. Canada became a haven for those who still believed in these ideas. Hinton was in Toronto, and one of LeCun's old colleagues from Bell Labs, Yoshua Bengio, another Paris-born researcher, oversaw a lab at the University of Montreal. During these years, Ian Goodfellow applied to graduate schools in computer science, and several offered him a spot, including Stanford, Berkeley, and Montreal. He preferred Montreal, but when he visited, a Montreal student tried to talk him out of it. Stanford was the number-three-ranked computer science program in North America. Berkeley was number four. And both were in sunny California. The University of Montreal was ranked somewhere around a hundred fifty, and it was cold.
“‘Stanford! One of the most prestigious universities in the world!’ this Montreal student told him as they walked through the city in late spring, snow still on the ground. ‘What the hell are you thinking?’
“‘I want to study neural networks,’ Goodfellow said.
“The irony was that as Goodfellow explored neural networks in Montreal, one of his old professors, Andrew Ng, after seeing the research that continued to emerge from Canada, was embracing the idea in his 'lab at Stanford. But he was very much an outlier, both at his own university and across the wider community, and he didn't have the data needed to convince those around him that neural networks were worth exploring. During these years, he gave a presentation at a workshop in Boston that trumpeted neural networks as the wave of the future. In the middle of his talk, the Berkeley professor Jitendra Malik, one of the de facto leaders of the computer vision community, stood up, Minsky-like, and told him this was nonsense, that he was making his self-satisfied claim with absolutely no evidence to support it.
“Around the same time, Hinton submitted a paper to NIPS, the conference where he would later auction off his company. This was a conference conceived in the late 1980s as a way for researchers to explore neural networks of all kinds, the biological as well as the artificial. But the conference organizers rejected Hinton's paper because they had accepted another neural network paper and thought it would be unseemly to accept two in the same year. ‘Neural’ was a bad word, even at a conference dedicated to Neural Information Processing Systems. Across the field, neural networks showed up in less than 5 percent of all published research papers. When submitting papers to conferences and journals, hoping to improve their chances of success, some researchers would replace the words ‘neural network’ with very different language, like ‘function approximation’ or ‘nonlinear regression.’ Yann LeCun removed the word ‘neural’ from the name of his most important invention. ‘Convolutional neural networks’ became ‘convolutional networks’
“Still, papers that LeCun viewed as undeniably important were rejected by the AI establishment, and when they were, he could be openly combative, adamant that his views were the right views. Some saw this as unfettered confidence. Others believed it betrayed an insecurity, an underlying remorse that his work wasn't recognized by the leaders of the field. One year, Clement Farabet, one of his PhD students, built a neural network that could analyze a video and separate different kinds of objects—the trees from the buildings, the cars from the people. It was a step toward computer vision for robots or self-driving cars, able to perform its task with fewer errors than other methods and at faster speeds, but reviewers at one of the leading vision conferences summarily rejected the paper. LeCun responded with a letter to the conference chair saying that the reviews were so ridiculous, he didn't know how to begin writing a rebuttal without insulting the reviewers. The conference chair posted the letter online for all to see, and though he removed LeCun's name, it was obvious who had written it.”





COMMENTS (0)
Notice: Trying to access array offset on value of type bool in /home/customer/www/delanceyplace.com/public_html/cmsAdmin/plugins/websiteComments/websiteComments.php on line 279