Skip to content
reflectionbeam

How to use

How to use Beam today

Beam is in a gradual beta. Here is every way to get at it right now, working code for the OpenAI-compatible API, and what changes when the open weights arrive.

Your options right now

From zero to the first request

  1. 01

    Join the waitlist

    Create an account on platform.reflection.ai. New sign-ups wait until their access is enabled.

  2. 02

    Create an API key

    In the platform, open API Keys and create a key. Store it in the REFLECTION_API_KEY environment variable.

  3. 03

    Point your client at Reflection

    Use the base URL https://api.reflection.ai/openai/v1 with any OpenAI SDK, and the model ID Beam-501B-A23B.

  4. 04

    Send a chat completion

    Call Chat Completions as usual. Optionally set reasoning_effort to trade speed for depth.

Code examples

The endpoint follows the OpenAI Chat Completions format, so the official OpenAI SDKs work after changing the base URL and key.

First request

curl https://api.reflection.ai/openai/v1/chat/completions \
  -H "Authorization: Bearer $REFLECTION_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Beam-501B-A23B",
    "messages": [{"role": "user", "content": "Explain Mixture-of-Experts in two sentences."}],
    "reasoning_effort": "medium"
  }'

Streaming

stream = client.chat.completions.create(
    model="Beam-501B-A23B",
    messages=[{"role": "user", "content": "Write a haiku about sparse experts."}],
    stream=True,
)

for chunk in stream:
    delta = chunk.choices[0].delta
    # reasoning arrives first in delta.reasoning_content, then the answer in delta.content
    if delta.content:
        print(delta.content, end="", flush=True)

Environment variables

export REFLECTION_API_KEY="<your API key>"

# Tools that read the standard OpenAI variables and use Chat Completions
export OPENAI_BASE_URL="https://api.reflection.ai/openai/v1"
export OPENAI_API_KEY="$REFLECTION_API_KEY"

# Check which models your key can use
curl https://api.reflection.ai/openai/v1/models -H "Authorization: Bearer $REFLECTION_API_KEY"

Reasoning effort

Beam always reasons; there is no way to switch it off. The reasoning_effort parameter controls how much. The default is medium.

ValueBehavior
lowLeast reasoning, lowest latency. For simple tasks.
mediumBalance of quality and latency. Used when you omit the parameter.
highMore reasoning for harder problems.
xhighMore reasoning than high.
maxThe most reasoning, for the hardest problems.

Reasoning tokens count toward max_completion_tokens and rate limits. If the limit is too low, the model can run out while still reasoning and return an empty answer with finish_reason "length". The reasoning text comes back separately in reasoning_content.

What works and what doesn't

Supported

  • POST /chat/completions, GET /models, GET /models/{model}
  • Streaming, including usage in the final chunk
  • Tool calling with tool_choice and parallel_tool_calls
  • Structured outputs: json_object and json_schema
  • temperature, top_p, frequency_penalty, presence_penalty, max_completion_tokens, seed
  • system, developer, user, assistant and tool roles

Not supported

  • Responses, Embeddings, Images, Audio, Files, Batch and Assistants APIs
  • Image, audio or file inputs: Beam is text-only
  • n other than 1, logprobs, logit_bias
  • Parameters such as user, metadata, audio or prediction (they return an error)
  • stop sequences are accepted but have no effect

Rate limits

Limits apply per organization, in requests and tokens per minute and per UTC day, plus a cap on concurrent requests. Exceeding one returns HTTP 429 with a Retry-After header. Daily limits reset at 00:00 UTC. Reflection has not yet published how to request higher limits, and there is no public price list for the beta.

Coding agents

Mirror CLI connects to Beam with nothing but an API key. Pi, OpenCode and Hermes work through their OpenAI-compatible provider settings: base URL https://api.reflection.ai/openai/v1, model Beam-501B-A23B, key from REFLECTION_API_KEY. Mirror does not support native Windows.

curl -fsSL https://raw.githubusercontent.com/reflection-oss/mirror-beta/main/install.sh | bash

cd your-project
mirror        # uses Reflection and Beam-501B-A23B by default

Running Beam yourself

When the weights ship, Beam can be served with open-source engines; Reflection says it is working on integrations and launch partners. Plan for a lot of memory: all 501 billion parameters must sit in GPU memory even though only 23 billion are active per token. That is roughly 1 TB at BF16, 500 GB at FP8 and 250 GB at 4-bit, before the KV cache.