Wednesday, September 30, 2026
HomeTechnologyMeta releases Llama 4, a new crop of flagship AI models

Meta releases Llama 4, a new crop of flagship AI models


Meta has released a new collection of AI models, Llama 4, in its Llama family โ€” on a Saturday, no less.

There are four new models in total: Llama 4 Scout, Llama 4 Maverick, and Llama 4 Behemoth. All were trained on โ€œlarge amounts of unlabeled text, image, and video dataโ€ to give them โ€œbroad visual understanding,โ€ Meta says.

The success of open models from Chinese AI lab DeepSeek, which perform on par or better than Metaโ€™s previous flagship Llama models, reportedly kicked Llama development into overdrive. Meta is said to have scrambled war rooms to decipher how DeepSeek lowered the cost of running and deploying models like R1 and V3.

Scout and Maverick are openly available on Llama.com and from Metaโ€™s partners, including the AI dev platform Hugging Face, while Behemoth is still in training. Meta says that Meta AI, its AI-powered assistant across apps including WhatsApp, Messenger, and Instagram, has been updated to use Llama 4 in 40 countries. Multimodal features are limited to the U.S. in English for now.

Some developers may take issue with the Llama 4 license.

Users and companies โ€œdomiciledโ€ or with a โ€œprincipal place of businessโ€ in the EU are prohibited from using or distributing the models, likely the result of governance requirements imposed by the regionโ€™s AI and data privacy laws. (In the past, Meta has decried these laws as overly burdensome.) In addition, as with previous Llama releases, companies with more than 700 million monthly active users must request a special license from Meta, which Meta can grant or deny at its sole discretion.

โ€œThese Llama 4 models mark the beginning of a new era for the Llama ecosystem,โ€ Meta wrote in a blog post. โ€œThis is just the beginning for the Llama 4 collection.โ€

Image Credits:Meta

Meta says that Llama 4 is its first cohort of models to use a mixture of experts (MoE) architecture, which is more computationally efficient for training and answering queries. MoE architectures basically break down data processing tasks into subtasks and then delegate them to smaller, specialized โ€œexpertโ€ models.ย 

Maverick, for example, has 400 billion total parameters, but only 17 billion active parameters across 128 โ€œexperts.โ€ (Parameters roughly correspond to a modelโ€™s problem-solving skills.) Scout has 17 billion active parameters, 16 experts, and 109 billion total parameters.

According to Metaโ€™s internal testing, Maverick, which the company says is best for โ€œgeneral assistant and chatโ€ use cases like creative writing, exceeds models such as OpenAIโ€™s GPT-4o and Googleโ€™s Gemini 2.0 on certain coding, reasoning, multilingual, long-context, and image benchmarks. However, Maverick doesnโ€™t quite measure up to more capable recent models like Googleโ€™s Gemini 2.5 Pro, Anthropicโ€™s Claude 3.7 Sonnet, and OpenAIโ€™s GPT-4.5.

Scoutโ€™s strengths lie in tasks like document summarization and reasoning over large codebases. Uniquely, it has a very large context window: 10 million tokens. (โ€œTokensโ€ represent bits of raw text โ€” e.g. the word โ€œfantasticโ€ split into โ€œfan,โ€ โ€œtasโ€ and โ€œtic.โ€) In plain English, Scout can take in images and up to millions of words, allowing it to process and work with extremely lengthy documents.

Scout can run on a single Nvidia H100 GPU, while Maverick requires an Nvidia H100 DGX system or equivalent, according to Metaโ€™s calculations.

Metaโ€™s unreleased Behemoth will need even beefier hardware. According to the company, Behemoth has 288 billion active parameters, 16 experts, and nearly two trillion total parameters. Metaโ€™s internal benchmarking has Behemoth outperforming GPT-4.5, Claude 3.7 Sonnet, and Gemini 2.0 Pro (but not 2.5 Pro) on several evaluations measuring STEM skills like math problem solving.

Of note, none of the Llama 4 models is a proper โ€œreasoningโ€ model along the lines of OpenAIโ€™s o1 and o3-mini. Reasoning models fact-check their answers and generally respond to questions more reliably, but as a consequence take longer than traditional, โ€œnon-reasoningโ€ models to deliver answers.

Meta Llama 4
Image Credits:Meta

Interestingly, Meta says that it tuned all of its Llama 4 models to refuse to answer โ€œcontentiousโ€ questions less often. According to the company, Llama 4 responds to โ€œdebatedโ€ political and social topics that the previous crop of Llama models wouldnโ€™t. In addition, the company says, Llama 4 is โ€œdramatically more balancedโ€ with which prompts it flat-out wonโ€™t entertain.

โ€œ[Y]ou can count on [Lllama 4] to provide helpful, factual responses without judgment,โ€ a Meta spokesperson told TechCrunch. โ€œ[W]eโ€™re continuing to make Llama more responsive so that it answers more questions, can respond to a variety of different viewpoints [โ€ฆ] and doesnโ€™t favor some views over others.โ€

Those tweaks come as some White House allies accuse AI chatbots of being too politically โ€œwoke.โ€

Many of President Donald Trumpโ€™s close confidants, including billionaire Elon Musk and crypto and AI โ€œczarโ€ David Sacks, have alleged that popular AI chatbotsย censor conservative views. Sacks has historicallyย singled outย OpenAIโ€™s ChatGPT as โ€œprogrammed to be wokeโ€ and untruthful about political subject matter.

In actuality, bias in AI is an intractable technical problem. Muskโ€™s own AI company, xAI, hasย struggledย to create a chatbot that doesnโ€™t endorse some political views over others.

That hasnโ€™t stopped companies including OpenAI from adjusting their AI models to answer more questions than they would have previously, in particular questions relating to controversial subjects.



Source link

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular

Recent Comments

Translate ยป