51¶ŻÂț

If only they could think

Should the Turing test be blamed for the ills that beset artificial intelligence? Or is it still the ultimate goal of every aspiring computer?

CAN A science be based on winning a parlour game? It is a curious-sounding game, that on the face of it involves nothing more than a man and a woman sending messages on computer screens to mislead a judge into thinking they are not the sex they appear to be.

Clearly the answer is yes, if some recent critics of artificial intelligence are to be believed. Only now some AI researchers are rebelling. They want to stop playing the game, which was bequeathed to them by the founder of their discipline as the premise for assessing machine intelligence. But will they ever be able to shake off the founder’s influence?

It was in 1950 that British mathematician and computer pioneer Alan Turing published a paper entitled “Computing Machinery and Intelligence” in the philosophical journal Mind with a description of his “imitation game”. Turing chose Mind because then, even when there were only a handful of computers on the planet, he foresaw that the machines would be dogged by the question of whether or not they were capable of “intelligence.”

Turing dreamt up the imitation game as a way to head off such debate. The game consists of a number of written conversations between man, woman and judge (sex unspecified). The man’s goal is to pretend to be a woman, the woman’s to resist being thought of as a man; and the judge’s to decide which is the woman.

If a computer, said Turing, could take the place of the man, and convince the judge that it is a woman (not a machine) 70 per cent of the time, that would constitute a working definition of the computer being able to “think”. Although some sources suggest that the “gender” issue was a red herring, the basic principle of a computer being able to convince a human of its humanity remains the same. Or as AI pioneer John McCarthy of Stanford University says, “Turing regarded it [this deception] as a sufficient condition. In arguing with philosophers who claim a machine can’t be intelligent, he could say ‘Do you mean that if its behaviour couldn’t be distinguished from that of a human, you still wouldn’t regard it as intelligent? Then we have nothing left to discuss’.”

Sceptical philosophers have had a lot to discuss in the intervening forty years. AI has acquired a reputation as a science based on empty boasts and capable of only weak experimental results. Frederick Allen, writing in the American magazine The Atlantic last year, spoke for many critics when he claimed that: “Today, traditional AI is a backwater at best, and the confidence with which it was once pursued seems unimaginable. Nobody has ever designed a program that can converse at all convincingly on a single subject, and the field has splintered into disparate parts 
 The grand vision has nearly vanished.” The LA Times columnist Michael Schrage cruelly jibed recently that thanks to “an impressive array of pompously wrong predictions, America’s artificial intelligentsia has a hard-earned reputation as the arrogant booboisie of bit-twiddlers”. In some computer science departments, they joke that “If it works, it’s no longer AI”, adds Schrage.

As funding threatens to dry up, some AI researchers are now trying to reinvigorate the field by reassessing Turing’s strategy for silencing those troublesome philosophers. “We must explicitly reject the Turing test [as the game is now called] in order to find a more mature description of our goals -it is time to move it from the textbooks to the history books,” say Pat Hayes of the University of Illinois’s Beckman Institute, and Kenneth Ford of the University of West Florida.

Mechanical transvestite

Hayes and Ford used the 14th joint conference on AI, held in Montreal last year, to slam the test as “harmful” to AI, “damaging its public reputation and its own intellectual coherence”. They allege that Turing’s experimental design is an attempt to confirm the “null hypothesis” – that is, trying to find nothing.

They also express concern about the emphasis on gender in the original game. A clever judge would be looking out for signs of femininity in the answers of the contestants. Women, for example, tend to be more specific about colours, distinguishing scarlet and crimson, while men tend to simply say red. “The gender test is not a test of making an artificial human but a mechanical transvestite,” snipes Hayes.

A program playing the Turing test, says Hayes, is not so much trying to display comprehension of language, as wield it with a ridiculous level of sophistication. “It would need more than ordinary conversational abilities, it would have to be an expert on making a good impression. It would have to be clever, to lie, to cheat, and dissemble. So we must not make an artificial intelligence but a mechanical trickster – an artificial con-artist.”

A way to judge the basis of Hayes and Ford’s dismay is to look at how today’s computers are trying to pass the test. In 1991, a New York-based computer consultant called Hugh Loebner provided the financial backing for competitions bearing his name in which human judges would sit and type at computer terminals with unseen computer and human entrants. The aim in the first three such competitions, held in Boston, was to “fool” 10 human judges in the course of three hours of typed interchanges about which was human and which not.

Loebner promised $100 000 for a program capable of doing this, but the entrants were so weak that it never looked for a moment as if he would have to get out his chequebook. Instead, in past years a $2000 prize has been awarded for the program judged to be “most human”. Most of these have succeeded by being programmed to stick to defined topics of conversation such as the Rolling Stones, sex, pets, or the O. J. Simpson trial – and even then they don’t do very well:

Topic: Classic Star Trek

Judge: Who are the Klingons?

Program: I don’t know who you’re talking about.

Judge: You’re the trekker.

Program: Harry Mudd was one of the more comical characters.

Judge: Harry Mudd? Who was Harry Mudd?

Program: No I haven’t.

Judge: Me neither.

Program: Okay, I will.

Judge: Why did you choose to converse about Star Trek?

Program: Their minds were stored in these white globes.

Judge: Whose minds?

Program: Spock’s ears are pointed and there was this strange hippy guy in “The way to Eden” that had these bizarre pretzel-shaped ears.

Judge: So how big were the ears?

Program: There must have been 70 or 80. I’m not sure.

Judge: Ears?

Program: I’m sure I would know if I could see it on TV.

Often the best results in Loebner competitions are achieved by tricks such as deliberately mistyping answers, creating the illusion of “listening” by repeating snatches of a user’s input back (like the famous 1960s’ AI therapist program Eliza – see “Oh look, he’s brought me a present”, 51¶ŻÂț, 16 September 1995), steering the conversation the computer’s way when it is in trouble by introducing controversial statements or throwing in humour to seem more human. As one team from the Centre for Machine Translation at Carnegie Mellon University in Pittsburgh, Pennsylvania, admitted in a background document on its system: “Here we unashamedly describe some of the better tricks, confident in the belief that when someday a computer program does pass the test, it will use many of them”.

Ironically, one contestant’s efforts were “rejected” as not human because its answers on literature were too detailed. The female Shakespeare scholar in question was dismayed. As Hayes quips: “The ability to produce paragraphs of well-written English is plainly now considered an inhuman ability.”

The Loebner test may be a bit of fun, but researchers who have invested a large part of their lives in search of an intelligent machine appear to fair no better. “Expert systems” – AI programs that try to consolidate enough expertise on a defined area of knowledge, such as areas of medicine or process control – are incapable of any interaction beyond their tightly-defined domains. Even the best ones, such as MYCIN, an expert system programmed to answer questions about eye diseases, do not know what time is, and are much less able to negotiate their way in the real world than infants.

Hayes and Ford believe that these limitations, along with the well-publicised failings of the Loebner competition have contributed to giving AI a bad name. “Last year, a couple in a bar asked me what I did for a living. With some trepidation, I replied that I worked in AI. ‘AI,’ said the lady, ‘I think I’ve read about that. They tried it, didn’t they, last year up in Boston?’ ‘Yes,’ said her husband. ‘But I thought that the machine failed the test. So why are you still working on it?'” jokes Hayes.

Why indeed? And it is not just remarks in bars that are forcing researchers to reassess their work. One of AI’s most bitter critics is John Searle, professor of philosophy at the University of California at Berkeley. Searle scored a near-mortal blow to AI in 1980 with his insistence in a paper called “Minds, Brains and Programs” that at best all it would ever achieve is a superficial likeness of intelligent behaviour.

Probably his most famous stab at AI is the Chinese room argument. Searle contends that machine intelligence is actually like passing English symbols into a sealed chamber. A person is hiding inside the room who simply looks up the given character in a translation table, then passes out a Chinese hexagram equivalent. In other words, computers do not understand what they are passing around, even if from outside it looks as if what emanates from the Chinese room is a perfectly intelligent translation.

In this light, even if a computer successfully played the imitation game, Searle would still say that it did not display intelligence. “Formal symbol-manipulations [of the kind computers do] by themselves don’t have any intentionality, they are quite meaningless,” he says. “They have only a syntax but no semantics. Such intentionality as computers appear to have is solely in the minds of those who program them and those who use them, those who send in the input and those who interpret the output.”

But many computer scientists resist Searle’s attack. They are driven by a grand vision of building intelligent devices. Even critics of the Turing test such as Hayes and Ford are still in search of the intelligent machine. But how will they know when they have developed one? For starters, they will not be creating an alternative Turing test. “Having such a neatly-defined aim is the sign of an immature discipline. What is the Turing test for civil engineering or physics?”

Five hundred years ago, alchemists sought to pass their own Turing test by searching for the philosopher’s stone, a magical stone or substance that could transmute base metal into gold. “But they only became chemists after they stopped looking for that stone. As a call to arms, Turing’s paper was a magical success. But that was 1950. It’s time to put alchemy behind us,” says Hayes.

The historical analogy is chosen with care. Hayes and Ford think AI must be freed from the obligation to copy nature before it can find its true path. “Natural systems – products of evolution in a complex world – tend to be very intricate, ad hoc devices. We need to know what we’re looking for when we examine them, and direct imitation isn’t a good way to discover what it is we’re looking for,” says Hayes.

People consider a transatlantic trip on Concorde “flying”, but birds can’t fly at Mach 3 nor can Jumbo jets catch salmon in a stream or land in trees. Thus “artificial flight” and “artificial intelligence” are both misnomers, says Hayes. AI should not be about copying human intelligence (as the Turing test implies) but more about cognition in terms of computation. That means understanding that which can be deemed intelligent – whether it is exhibited by ants, humans or machines. In other words, the critics see AI transmuting into the science of computation, rather as alchemy became the science of chemistry.

“As we develop a general science of computation, it is the aspects of thought which are ‘not’ distinctively human that seem the most fundamental,” says Ford. So linguists, psychologists, cognitive scientists and some philosophers are beginning to move toward the position that being human is not a sine qua non of intelligence, and that there are core computational functions that are useful – being able to tell red from blue, for example, or provide accurate information on what happens if a patient takes two different kinds of drug – and which can be reproduced by writing programs.

There are many AI scientists who reject Searle, but who also disagree with Hayes and Ford in their condemnation of the Turing test. Danny Bobrow, developer of some pioneering AI systems and now a researcher at Xerox PARC in Palo Alto, argues that, “the Turing test is a systems test – AI programs aren’t really worth it of they can’t all be put together”. It is all very well building systems that can cope with aspects of intelligent behaviour, such as understanding speech or discovering mathematical theorems, but we will only have intelligence as we understand it when we can put it all together.

Bobrow uses the old Indian philosopher’s joke, in which three blind wise men each grab disparate parts of something and fail to synthesise their findings. “You don’t know if you’ve got a snake, a tree trunk and a big wall or if you’ve in fact got something that proves to be an elephant.” So unless we synthesise AI’s findings into a suite of programs that could pass the test, we don’t know if we’ve actually cracked the problem – so we’re back to Turing and his test.

Intelligent species

Others feel that Hayes and Ford are wrong to blame the test for all AI’s ills. “I don’t think that even Turing regarded the imitation game as a definition,” says McCarthy. “His other writings on machine intelligence don’t even mention it. Anyway, I never regarded the imitation game as a useful goal.”

Still others want to stick with the test, confident that it is the benchmark of what an artificial – and they think “higher” – intelligence will achieve. One of the scientists behind the Loebner competition is Robert Epstein, of the Cambridge Center for Behavioral Studies in California, who dismisses Hayes’s and Ford’s objections as “very primitive,” claiming, “the major problem is taking Turing too literally. He never pretended to give the details of a practical test”.

Epstein is an AI visionary of the old school. “I suspect that much more is coming – that the winning program will be the first member of a new intelligent species that will populate our planet virtually overnight, by means of replicating over the Internet,” he says. “We’re going to have to adapt, just as we would if intelligent aliens settled here.”

Such optimism – or daftness – is a poor ad for what AI researchers have achieved in problem solving, natural speech understanding, vision, robotics and expert systems. Yet it may be the curse of AI that the Turing test still defines what it is they can and cannot do.

A fitting footnote to the controversies over what the Turing test is and isn’t is the fact that the Loebner prize has been cancelled. The prize committee is now set to launch an alternative competition which will be designed in “the spirit of Turing’s original proposal”.

The controversy over what exactly that is, and what we should make of it fifty years on, is set to continue. In the meantime, Al in the sense of the science fiction author’s dream of creating an artificial intelligence that will be a kind of mechanical sibling to us, seems as far away as ever.