Claude 2 Brings a Longer Context Window to AI Chat

Anthropic launched Claude 2 as a chatbot competitor to ChatGPT, with claimed improvements in conversation, coding, math, and safety. Its ability to process up to 75,000 words (100,000 tokens) at a time is a key feature, though the company says the model can still produce misinformation.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The article describes a routine chatbot launch with broader capabilities and acknowledges misinformation risk, but no clear societal harm or degradation.

Claude 2 Brings a Longer Context Window to AI Chat

Anthropic has launched Claude 2, a chatbot positioned to compete with ChatGPT, Google Bard, and Bing Chat. The company says the updated model improves on its earlier system in several areas, while its unusually large context window could help it work with longer documents and more information at once.

More text in view at one time

Claude 2 can process up to 75,000 words (100,000 tokens) at a time, according to the article. That is substantially more than the standard ChatGPT limit of 3,000 words cited in the report. A larger context window gives a chatbot more of a document or conversation to consider while preparing a response.

That capacity could be useful for tasks involving long material. Anthropic says Claude 2 can help write documents, memos, letters, stories, technical documentation, and books. With more context available, a user may be able to ask about a larger body of text without dividing it into as many separate pieces.

The expanded window may also support a wider variety of tasks and improve response quality by letting the model draw on more of the supplied material. Anthropic had announced the extra-large context window for its first Claude model in May, the article says.

Anthropic reports gains in tests

The company says Claude 2 has better conversational skills, offers clearer explanations of its reasoning, and produces more harmless output. It also reports improvements in programming, math, and thinking skills. These are company claims, and the article presents benchmark results as one way to assess them.

On the multiple-choice section of the American Bar Exam, Claude 2 scored 76.5 percent, matching GPT-4 in the comparison described. GPT-3.5, identified in the article as the free ChatGPT, averages about 50 percent. On the Codex HumanEval Python programming test, Claude 2 scored 71.2 percent, compared with 56.0 percent for Claude 1.3.

For GSM8k elementary school math problems, the reported score was 88.0 percent, up from 85.2 percent for Claude 1.3. Those results point to improvements on the tests listed, but they do not establish how the chatbot will perform on every real-world task. The article does not present benchmark scores as proof that the model is error-free.

Safety remains a work in progress

Anthropic says Claude 2 took about two months to develop. About 35 people worked directly on the AI model, with another 150 in supporting roles. The company says it paid particular attention to safety during development.

Anthropic uses an AI-based feedback mechanism to evaluate AI-generated content and optimize the model, rather than involving humans in that evaluation, according to the article. It also sets ground rules through a kind of constitution based on Apple's T&Cs, among other guidelines.

In red team testing, where testers deliberately try to provoke mistakes, Anthropic says Claude 2 achieved twice the pleasant user experience of its predecessor. The company also acknowledges that Claude 2 can still hallucinate or provide misinformation, and says many hurdles remain. Those limits matter for anyone relying on chatbot responses: stronger test results do not remove the need to assess what the system produces.

Availability and business use

Claude 2's web chatbot launched as a free beta in the US and UK. Anthropic says the Claude 2 API is available to business customers for the same price as Claude 1.3. The company says thousands of businesses are already using the API.

The article names Jasper, a platform for generative AI marketing copy, and Sourcegraph, a code AI platform, among Anthropic's partners. Sourcegraph uses Claude's enhanced reasoning capabilities and larger context windows to help developers write, fix, and maintain code, according to the report.

Anthropic says additional capabilities will arrive gradually over the coming months. For now, Claude 2's launch combines reported benchmark gains with a larger capacity for handling text, while its own maker cautions that misinformation and other challenges have not been solved.