AI Inference: Meta Teams with Cerebras on Llama API

Sunnyvale, CA — Meta has teamed with Cerebras on AI inference in Meta’s new Llama API, combining Meta’s open-source Llama fashions with inference expertise from Cerebras.

Builders constructing on the Llama 4 Cerebras mannequin within the API can anticipate speeds as much as 18 occasions sooner than conventional GPU-based options, in response to Cerebras. “This acceleration unlocks a wholly new technology of purposes which are inconceivable to construct on different expertise. Conversational low latency voice, interactive code technology, on the spot multi-step reasoning, and real-time brokers — all of which require chaining a number of LLM calls — can now be accomplished in seconds fairly than minutes,” Cerebras stated.

By partnering with Meta to serve Llama fashions from Meta’s new API service, Cerebras positive factors publicity to an expanded developer viewers and deepens its enterprise and partnership with Meta and their unimaginable groups.

Since launching its inference options in 2024, Cerebras has delivered the world’s quickest Llama inference, serving billions of tokens by way of its personal AI infrastructure. The broad developer group now has direct entry to a sturdy, OpenAI-class different for constructing clever, real-time methods — backed by Cerebras pace and scale.

“Cerebras is proud to make Llama API the quickest inference API on the earth,” stated Andrew Feldman, CEO and co-founder of Cerebras. “Builders constructing agentic and real-time apps want pace. With Cerebras on Llama API, they will construct AI methods which are essentially out of attain for main GPU-based inference clouds.”

Cerebras is the quickest AI inference resolution as measured by third occasion benchmarking web site Synthetic Evaluation, reaching over 2,600 token/s for Llama 4 Scout in comparison with ChatGPT at ~130 tokens/sec and DeepSeek at ~25 tokens/sec.

Builders will have the ability to entry to the quickest Llama 4 inference by choosing Cerebras from the mannequin choices inside the Llama API. This streamlined expertise will make it straightforward to prototype, construct, and scale real-time AI purposes. To join early entry to the Llama API and to expertise Cerebras pace at present, go to www.cerebras.ai/inference.

Source link

Why Teams Rely on Data Structures

Definite Raises $10M for AI-Native Data Stack

EdgeConneX and Lambda to Build AI Factory Infrastructure in Chicago and Atlanta

Elon Musk and X reach settlement with axed Twitter workers

I Tried Buying a Car Through Amazon: Here Are the Pros, Cons

Amazon and eBay to pay ‘fair share’ for e-waste recycling

Artificial Intelligence Concerns & Predictions For 2025

Barbara Corcoran: Entrepreneurs Must ‘Embrace Change’

Most Popular

Djdhdj

This City Is the Best Place to Be an Entrepreneur Right Now

Student debit crises | by Myaseen | Dec, 2024

Our Picks

Elon Musk and X reach settlement with axed Twitter workers

Labubu Could Reach $1B in Sales, According to Pop Mart CEO

Unfiltered Roleplay AI Chatbots with Pictures – My Top Picks

AI Inference: Meta Teams with Cerebras on Llama API

Related Posts