1. You Are At:
  2. English News
  3. Technology
  4. ChatGPT- What exactly is it? How does it work?

ChatGPT- What exactly is it? How does it work?

ChatGPT has been trained for using 'Reinforcement Learning from Human Feedback' (RLHF), by using the same methods as InstructGPT, but with some differences in the data collection setup. Here is everything you need to know about the AI-driven answering bot.

Saumya Nigam Edited By: Saumya Nigam @snigam04 Noida Updated on: January 18, 2023 17:37 IST
Image Source : OPENAI Openai

As the world goes gaga over artificial intelligence (AI)-driven chatbot called ChatGPT which is capable to write poems and essays and could also make humorous comments like your friendly buddy. The new conversational AI has opened several frontiers for real-life use cases shortly which could only be handled with care.

According to OpenAI, the founding company of chatGPT, they have trained an AI model which could interact conversationally.

The dialogue format makes it possible for ChatGPT to answer follow-up questions, admit its mistakes, challenge incorrect premises, and reject inappropriate requests.

ChatGPT is a sibling model to "InstructGPT", which is trained to follow instructions in a prompt and provide a detailed response, according to OpenAI which was acquired by Microsoft for $1 billion.

How does it work? 

The company has trained the model using 'Reinforcement Learning from Human Feedback' (RLHF), using the same methods as InstructGPT, but with slight differences in the data collection setup.

OpenAI has said, "We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides - the user and an AI assistant."

The teams further have given the trainers access to model-written suggestions to help them compose their responses.

India Tv - ChatGPT

Image Source : OPENAIChatGPT

"We mixed this new dialogue dataset with the InstructGPT dataset, which we transformed into a dialogue format," the company stated.

To create a reward model for reinforcement learning, it took conversations that AI trainers had with the chatbot.

"We randomly selected a model-written message, sampled several alternative completions, and had AI trainers rank them. Using these reward models, we can fine-tune the model using 'Proximal Policy Optimisation'. We performed several iterations of this process," explained OpenAI.

What are the limitations of ChatGPT?

ChatGPT sometimes writes plausible-sounding but incorrect or nonsensical answers at times. According to the company, fixing this issue is challenging, as during RL training, there's currently no source of truth and training the model to be more cautious causes it to decline questions that it can answer correctly.

India Tv - ChatGPT

Image Source : OPENAIChatGPT

Also, supervised training misleads the model because the "ideal answer depends on what the model knows, rather than what the human demonstrator knows".

ChatGPT is sensitive to tweaks to the input phrasing or attempting the same prompt multiple times. For example, given one phrasing of a question, the model can claim to not know the answer, but given a slight rephrase, can answer correctly, according to OpenAI.

The model is often excessively verbose and overuses certain phrases, such as restating that it's a language model trained by OpenAI.

"These issues arise from biases in the training data (trainers prefer longer answers that look more comprehensive) and well-known over-optimisation issues," the company admitted.

India Tv - ChatGPT

Image Source : OPENAIChatGPT

"While we've made efforts to make the model refuse inappropriate requests, it will sometimes respond to harmful instructions or exhibit biased behaviour. We're using the Moderation API to warn or block certain types of unsafe content, but we expect it to have some false negatives and positives for now," it added.

The company is currently collecting user feedback.

Inputs from IANS

Latest Technology News

Read all the Breaking News Live on indiatvnews.com and Get Latest English News & Updates from Technology