惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

D
Docker
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - Franky
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
A
About on SuperTechFans
博客园 - 【当耐特】
Microsoft Security Blog
Microsoft Security Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
The GitHub Blog
The GitHub Blog
雷峰网
雷峰网
博客园_首页
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
IT之家
IT之家
博客园 - 叶小钗
Google DeepMind News
Google DeepMind News
aimingoo的专栏
aimingoo的专栏
博客园 - 聂微东
B
Blog RSS Feed
H
Help Net Security
Recent Announcements
Recent Announcements
阮一峰的网络日志
阮一峰的网络日志
D
DataBreaches.Net
L
LangChain Blog
Vercel News
Vercel News

Replicate's blog

How to make remarkable videos with Seedance 2.0 – Replicate blog How to prompt Seedream 5.0 – Replicate blog Recraft V4: image generation with design taste – Replicate blog Run Isaac 0.1 on Replicate – Replicate blog Run FLUX.2 on Replicate – Replicate blog How to prompt Nano Banana Pro – Replicate blog Retro Diffusion's pixel art models are now on Replicate – Replicate blog Replicate is joining Cloudflare – Replicate blog Extract text from documents and images with Datalab Marker and OCR – Replicate blog How to prompt Veo 3.1 – Replicate blog IBM's Granite 4.0 is now on Replicate – Replicate blog Which image editing model should I use? – Replicate blog Introducing our new search API – Replicate blog Torch compile caching for inference speed – Replicate blog Announcing Replicate's remote MCP server – Replicate blog How to prompt Veo 3 with images – Replicate blog Open source video is back – Replicate blog Generate consistent characters – Replicate blog Bria is now on Replicate – Replicate blog How we optimized FLUX.1 Kontext [dev] – Replicate blog Compare AI video models – Replicate blog The FLUX.1 Kontext hackathon – Replicate blog How to prompt Veo 3 for the best results – Replicate blog Get the most from Google Veo 3 – Replicate blog FLUX.1 Kontext from the community – Replicate blog Use FLUX.1 Kontext to edit images with words – Replicate blog Generate incredible images with Google's Imagen 4 – Replicate blog Run OpenAI’s latest models on Replicate – Replicate blog NVIDIA H100 GPUs are here – Replicate blog Run 30,000+ LoRAs on Hugging Face with Replicate – Replicate blog
Run Meta Llama 3.1 405B with an API – Replicate blog
2024-07-23 · via Replicate's blog

Llama 3.1 is the latest language model from Meta. It features a massive 405 billion parameter model that rivals GPT-4 in quality, with a context window of 8000 tokens.

With Replicate, you can run Llama 3.1 in the cloud with one line of code.


Try Llama 3.1 in our API playground

Before you dive in, try Llama 3.1 in our API playground.

Try tweaking the prompt and see how Llama 3.1 responds. Most models on Replicate have an interactive API playground like this, available on the model page: https://replicate.com/meta/meta-llama-3.1-405b-instruct

The API playground is a great way to get a feel for what a model can do, and provides copyable code snippets in a variety of languages to help you get started.

Running Llama 3.1 with JavaScript

You can run Llama 3.1 with our official JavaScript client:

Install Replicate’s Node.js client library

Set the REPLICATE_API_TOKEN environment variable

(You can generate an API token in your account. Keep it to yourself.)

Import and set up the client

Run meta/meta-llama-3.1-405b-instruct using Replicate’s API. Check out the model’s schema for an overview of inputs and outputs.

To learn more, take a look at the guide on getting started with Node.js.

Running Llama 3.1 with Python

You can run Llama 3.1 with our official Python client:

Install Replicate’s Python client library

Set the REPLICATE_API_TOKEN environment variable

(You can generate an API token in your account. Keep it to yourself.)

Import the client

Run meta/meta-llama-3.1-405b-instruct using Replicate’s API. Check out the model’s schema for an overview of inputs and outputs.

To learn more, take a look at the guide on getting started with Python.

Running Llama 3.1 with cURL

Your can call the HTTP API directly with tools like cURL:

Set the REPLICATE_API_TOKEN environment variable

(You can generate an API token in your account. Keep it to yourself.)

Run meta/meta-llama-3.1-405b-instruct using Replicate’s API. Check out the model’s schema for an overview of inputs and outputs.

To learn more, take a look at Replicate’s HTTP API reference docs.

You can also run Llama using other Replicate client libraries for Go, Swift, and others.

About Llama 3.1 405B

Llama 3.1 405B is currently the only variant available on Replicate. This model represents the cutting edge of open-source language models:

  • 405 billion parameters: This massive model size allows for unprecedented capabilities in an open-source model.
  • Instruction-tuned: Optimized for chat and instruction-following tasks.
  • GPT-4 level quality: In many benchmarks, Llama 3.1 405B approaches or matches the performance of GPT-4.
  • Multilingual support: Trained on 8 languages including English, German, French, Italian, Portuguese, Hindi, Spanish, and Thai.
  • Extensive training: Trained on over 15 trillion tokens of data.

Responsible AI and Safety

Llama 3.1 comes with a strong focus on responsible AI development. Meta has introduced several tools and resources to help developers use the model safely and ethically:

  • Purple Llama: An open-source project that includes safety tools and evaluations for generative AI models.
  • Llama Guard 3: An updated input/output safety model.
  • Code Shield: A tool to help prevent unsafe code generation.
  • Responsible Use Guide: Guidelines for ethical use of the model.

We recommend reviewing these resources when building applications with Llama 3.1. For more information, check out the Purple Llama GitHub repository.

Example chat app

If you want a place to start, we’ve built a demo chat app in Next.js that can be deployed on Vercel:

Llama chat

Try it out on llama3.replicate.dev. Take a look at the GitHub README to learn how to customize and deploy it.

Keep up to speed

Happy hacking! 🦙