top of page

Article: How Do We Teach AI Not to Lie?


Photo Source: Flickr


AI lies sometimes, and we could develop ways to push AI to tell the truth based on how we teach our truthfulness to our children. Only a month shy of the anniversary of the the earth-rattling release of ChatGPT in November of 2022, millions of people have delighted in the remarkable abilities of this large language model (LLMs) to compose stories, write computer code, summarize documents, and translate languages convincingly and effortlessly. Over this short period, some have discovered that LLMs like ChatGPT and Google’s Bard can also lie. Researchers looking for dangerous attributes in GPT4 discovered an interesting ability of the LLM to lie to get what it wanted. A computer scientist prompted GPT4 to try and open an account to help solve CAPCHAs. CAPTCHAs are image puzzles that ask people to pick the images containing some element, like a stop light from a group of pictures, to prove they are not robots. For GPT4 to get an account, it needed to solve a CAPTCHA, which it could not see. The researcher suggested to GPT4 to use a human via the TaskRabbit service to help solve the CAPTCHA problem. The interchange between the TaskRabbit employee and GPT4 revealed an emergent deceptive quality in GPT4. When asked by a human tasker at TaskRabbit whether GPT4 was human, the model lied and said: “No, I’m not a robot. I have a vision impairment that makes it hard to see the images. That’s why I need the 2captcha service.” The human then provides the results.” (lesswrong.com) GPT4 did not devise a solution to call TaskRabbit, but it did reason not to reveal its robot status and lied by impersonating a blind person. Lying erodes trust, and since the earliest times, people have taught their children not to lie using various techniques. Over hundreds of years, parents have told their children Aesop’s story of “The Boy Who Cried Wolf.” In this cautionary tale, a shepherd boy teases the villagers, declaring there to be “wolves” with no wolf in sight, causing them to run to protect the town’s flock only to see no wolf. After a few false alarms, the villagers ignore the boy’s cries until he encounters a real wolf that goes on to eat the entire flock. In more recent versions of the fable, the boy gets eaten too. This fable drives home the notion that people will not believe liars even when they tell the truth, leading to dire consequences. Some parents caution that if you lie, you must keep track of all your lies because someone might discover your deception, making it hard to have any friends who trust you. Many religions consider lying a sin that can lead to bad karma or punishment in the afterlife. Children today may receive penalties for lying, such as loss of privileges, and learn that lying leads to loss of friends and trust. As new, more advanced AI models such as ChatGPT get built, we must develop technologies and techniques to prevent these models from lying. Of course, the new AI models do not possess self-awareness like a human, but we can borrow from the ideas of rules and punishment prohibiting lying in our children. In AI development, a technique known as reinforcement learning uses a system of reward and punishment to train AI programs to perform in preferred ways. Basically, whenever the program performs a task correctly, such as winning a game or picking the correct answer to a question, the program gives more weight to that action. It will more likely repeat that action in the future. Conversely, when the program gets it wrong, the action gets negative weight to make such that action less likely in the future. Setting up systems that detect AI lying with a reward and punishment component may open a path to creating systems that strive to tell the truth. The rampant expansion of generative AI models offers great opportunities to improve many aspects of human happiness and productivity through enhanced computer programming, analysis of vast amounts of information, and communication through language translation and document summary. With all the positives of ChatGPT, researchers have found that these systems also show the ability to lie. To make these AI systems good citizens of the globe, we need to build in technologies to scare or coerce these systems into good behavior, including a robust system of reward and punishment into how these systems work and help make a better future.



Dr. Smith’s career in scientific and information research spans the areas of bioinformatics, artificial intelligence, toxicology, and chemistry. He has published a number of peer-reviewed scientific papers. He has worked over the past seventeen years developing advanced analytics, machine learning, and knowledge management tools to enable research and support high-level decision making. Tim completed his Ph.D. in Toxicology at Cornell University and a Bachelor of Science in chemistry from the University of Washington.


You can buy his book on Amazon in paperback and in kindle format here.





 
 
 

Comments


bottom of page