How to use
How to use Beam today
Beam is in a gradual beta. Here is every way to get at it right now, working code for the OpenAI-compatible API, and what changes when the open weights arrive.
Your options right now
Reflection beta API
Sign up on the Reflection platform and join the waitlist. Once access is enabled you can create API keys and call Beam over an OpenAI-compatible endpoint.
Join the waitlistMirror CLI
Reflection's own coding agent for the terminal, on macOS (Apple Silicon) and Linux. It uses Beam by default and needs the same API key.
Mirror quickstartChat in the browser
BeamAI.chat is an independent chat site that will switch to Beam as soon as it is publicly available. Until then it runs a comparable open model and says so clearly.
Open BeamAI.chatDownload the weights
Not possible yet. Reflection has promised the weights under Apache 2.0 later in October 2026. We track the release daily.
Release trackerFrom zero to the first request
- 01
Join the waitlist
Create an account on platform.reflection.ai. New sign-ups wait until their access is enabled.
- 02
Create an API key
In the platform, open API Keys and create a key. Store it in the REFLECTION_API_KEY environment variable.
- 03
Point your client at Reflection
Use the base URL https://api.reflection.ai/openai/v1 with any OpenAI SDK, and the model ID Beam-501B-A23B.
- 04
Send a chat completion
Call Chat Completions as usual. Optionally set reasoning_effort to trade speed for depth.
Code examples
The endpoint follows the OpenAI Chat Completions format, so the official OpenAI SDKs work after changing the base URL and key.
First request
curl https://api.reflection.ai/openai/v1/chat/completions \
-H "Authorization: Bearer $REFLECTION_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "Beam-501B-A23B",
"messages": [{"role": "user", "content": "Explain Mixture-of-Experts in two sentences."}],
"reasoning_effort": "medium"
}'Streaming
stream = client.chat.completions.create(
model="Beam-501B-A23B",
messages=[{"role": "user", "content": "Write a haiku about sparse experts."}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta
# reasoning arrives first in delta.reasoning_content, then the answer in delta.content
if delta.content:
print(delta.content, end="", flush=True)Environment variables
export REFLECTION_API_KEY="<your API key>"
# Tools that read the standard OpenAI variables and use Chat Completions
export OPENAI_BASE_URL="https://api.reflection.ai/openai/v1"
export OPENAI_API_KEY="$REFLECTION_API_KEY"
# Check which models your key can use
curl https://api.reflection.ai/openai/v1/models -H "Authorization: Bearer $REFLECTION_API_KEY"Reasoning effort
Beam always reasons; there is no way to switch it off. The reasoning_effort parameter controls how much. The default is medium.
| Value | Behavior |
|---|---|
| low | Least reasoning, lowest latency. For simple tasks. |
| medium | Balance of quality and latency. Used when you omit the parameter. |
| high | More reasoning for harder problems. |
| xhigh | More reasoning than high. |
| max | The most reasoning, for the hardest problems. |
Reasoning tokens count toward max_completion_tokens and rate limits. If the limit is too low, the model can run out while still reasoning and return an empty answer with finish_reason "length". The reasoning text comes back separately in reasoning_content.
What works and what doesn't
Supported
- POST /chat/completions, GET /models, GET /models/{model}
- Streaming, including usage in the final chunk
- Tool calling with tool_choice and parallel_tool_calls
- Structured outputs: json_object and json_schema
- temperature, top_p, frequency_penalty, presence_penalty, max_completion_tokens, seed
- system, developer, user, assistant and tool roles
Not supported
- Responses, Embeddings, Images, Audio, Files, Batch and Assistants APIs
- Image, audio or file inputs: Beam is text-only
- n other than 1, logprobs, logit_bias
- Parameters such as user, metadata, audio or prediction (they return an error)
- stop sequences are accepted but have no effect
Rate limits
Limits apply per organization, in requests and tokens per minute and per UTC day, plus a cap on concurrent requests. Exceeding one returns HTTP 429 with a Retry-After header. Daily limits reset at 00:00 UTC. Reflection has not yet published how to request higher limits, and there is no public price list for the beta.
Coding agents
Mirror CLI connects to Beam with nothing but an API key. Pi, OpenCode and Hermes work through their OpenAI-compatible provider settings: base URL https://api.reflection.ai/openai/v1, model Beam-501B-A23B, key from REFLECTION_API_KEY. Mirror does not support native Windows.
curl -fsSL https://raw.githubusercontent.com/reflection-oss/mirror-beta/main/install.sh | bash
cd your-project
mirror # uses Reflection and Beam-501B-A23B by defaultRunning Beam yourself
When the weights ship, Beam can be served with open-source engines; Reflection says it is working on integrations and launch partners. Plan for a lot of memory: all 501 billion parameters must sit in GPU memory even though only 23 billion are active per token. That is roughly 1 TB at BF16, 500 GB at FP8 and 250 GB at 4-bit, before the KV cache.