Skip to main content
Groq’s Language Processing Units (LPUs) deliver the fastest inference available. Llama 3.3 70B runs at hundreds of tokens per second - ideal for real-time chat and low-latency agentic workflows.

Supported Models

Setup

1

Get an API key

Sign up at console.groq.com. Free tier available.
2

Set the environment variable

3

Verify

Environment Variables

string
required
Your Groq API key. Format: gsk_...

Configuration Example

Model Aliases

Usage Examples

Notes

  • Groq is ranked 5th in auto-selection priority after Anthropic, OpenAI, Azure, and Google.
  • llama-3.1-8b-instant is one of the cheapest available models at $0.05/1M input tokens.
  • Groq has a generous free tier with rate limits per day.
  • API is OpenAI-compatible - endpoint: https://api.groq.com/openai/v1