The “Dangerous” OpenAI Text Generator Recreate by Two Researchers

On Thursday, a pair of master graduates from Computer Science rolled out an AI text generator based upon GPT-2, an Elon Musk-backed OpenAI program that the company withheld from public release citing concerns over the societal impact of it.

May 03, 2020 - 10:36
Updated: 8 days ago
102.9k
The “Dangerous” OpenAI Text Generator Recreate by Two Researchers

Quick answer: In 2019 two researchers, Aaron and Vanya, recreated OpenAI's GPT-2 language model using about $50,000 of free Google cloud computing credits. They wanted to show that anyone could build such software, and said it carried no risk to society yet. OpenAI had held GPT-2 back, citing fear of misuse.

Key takeaways

  • The 2019 story is about two researchers rebuilding GPT-2 with free cloud credit to argue that such models were not an immediate risk.
  • It shows that a model once called too dangerous to release can be reproduced by small teams.
  • Treat this as news history; it says little about today's AI models.

However, the 2 researchers, Aaron and Vanya believe that the software does not possess any risk to society — not yet. According to Wired, the duo wanted to prove that anyone can develop the software, regardless of their economical status.

In order to replicate GPT-2, the duo used $50,000 worth of free cloud computing from Google. The research graduates also fed millions of webpages to the ML software, gathered by digging up links shared on Reddit.

Just like OpneAI GPT-2, the newly created software analyzed the language patterns and could be used up for many tasks — Translation, chatbots, coming up with unprecedented answers and more. However, the foremost alarming concern among specialists has been the creation of synthetic text, consequently, Fake news.

David Luan, vice president of engineering at OpenAI once told Wired, “It could be that someone who has malicious intent would be able to generate high-quality fake news”. Owing to this and other dangers, the team decided to withhold the model. However, it did put out a research paper.

Previously, there have been iterations of GPT-2. In fact, few people have released language models online based upon the OpenAI software. Of course, they’re not using the first model that used “8 million web pages”, however, it still uses the previous versions. You can try it out yourself.

While they’re smart for playing around, they don’t appear to provide logical statements. Wired, who tested out the original GPT-2 and new model as well, writes, “Machine learning software picks up the statistical patterns of language, not a true understanding of the world.”

To take this further with guided labs and an instructor, see our practical LLM engineering labs.

Related reading

Frequently Asked Questions

GPT-2 is a language model released by OpenAI in 2019 that generates text from a prompt. It was an early large language model and a step towards the GPT-3 and GPT-4 families.

OpenAI first held back the full model because it worried the tool could produce fake news and spam at scale. It later released the full model after seeing little evidence of misuse.

They trained a similar model using roughly $50,000 of free cloud computing credits from Google, plus public information about the architecture and similar web-scale training data.

Small and mid-size models are within reach using open-source tools and rented GPUs. Training frontier models still needs very large budgets, though fine-tuning open models is affordable.

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0
Aayushi Sinha

With a passion for staying on the cutting edge of technology trends, I am dedicated to delivering content that not only informs but also inspires. Whether you need in-depth analysis pieces, informative guides, or thought-provoking opinion pieces, I craft content that resonates with tech enthusiasts and professionals alike.