Skip to main content
Agents encapsulate models, prompting, tools, and other configuration. They can be defined as globals, or at runtime. They use threads to contain a series of messages used along the way, whether those messages are from a user, another Agent / LLM, or elsewhere. A thread can have multiple Agents responding, or be used by a single Agent. Agentic workflows are built up by combining contextual prompting (threads, messages, tool responses, RAG, etc.) and dynamic routing via LLM tool calls, structured LLM outputs, or a myriad of other techniques via custom code.

Basic Agent definition

These examples use the Bijection AI Gateway. See Getting Started for installation and requirements.
See below for more configuration options. Everything except the name can be overridden at the call site when calling the LLM, and many features available on the agent can be used without an Agent, if this way of organizing the work is not needed for your use case.

Dynamic Agent definition

You can define an Agent at runtime, which is useful if you want to create an Agent for a specific context. This allows the LLM to call tools without requiring the LLM to always pass through full context to each tool call. It also allows dynamically choosing a model or other options for the Agent.

Generating text with an Agent

To generate a message, you provide a prompt (as a string or a list of messages) to be used as context to generate one or more messages via an LLM, using calls like agent.streamText or agent.generateObject. The arguments to generateText and others are the same as the AI SDK, except you don’t have to provide a model. By default it will use the agent’s language model. There are also extra arguments that are specific to the Agent component, such as the promptMessageId which we’ll see below. See the full list of AI SDK arguments here The message history will be provided by default as context from the given thread. See LLM Context for details on how to configuring the context provided. Note: authorizeThreadAccess referenced below is a function you would write to authenticate and authorize the user to access the thread. You can see an example implementation in threads.ts. See chat/basic.ts or chat/streaming.ts for live code examples.

Streaming text

Streaming text follows the same pattern as the approach below, but with a few differences, depending on the type of streaming you’re doing. See streaming for more details.

Basic approach (synchronous)

Note: best practice is to not rely on returning data from the action. Instead, query for the thread messages via the useThreadMessages hook and receive the new message automatically. See below.

Saving the prompt then generating response(s) asynchronously

While the above approach is simple, generating responses asynchronously provide a few benefits:
  • You can set up optimistic UI updates on mutations that are transactional, so the message will be shown optimistically on the client until the message is saved and present in your message query.
  • You can save the message in the same mutation (transaction) as other writes to the database. This message can then be used and re-used in an action with retries, without duplicating the prompt message in the history. If the promptMessageId is used for multiple generations, any previous responses will automatically be included as context, so the LLM can continue where it left off. See workflows for more details.
  • Thanks to the idempotent guarantees of mutations, the client can safely retry mutations for days until they run exactly once. Actions can transiently fail.
Any clients listing the messages will automatically get the new messages as they are created asynchronously. To generate responses asynchronously, you need to first save the message, then pass the messageId as promptMessageId to generate / stream text.
Note that the action doesn’t need to return anything. All messages are saved by default, so any client subscribed to the thread messages will receive the new message as it is generated asynchronously.

Generating an object

Similar to the AI SDK, you can generate or stream an object. The same arguments apply, except you don’t have to provide a model. It will use the agent’s default language model.
Unfortunately, object generation doesn’t support using tools. One, however, is to structure your object as arguments to a tool call that returns the object. You can use a custom stopWhen to stop the generation when the tool call produces the result and use toolChoice: "required" to prevent the LLM from returning a text response.

Customizing the agent

The agent by default only needs a languageModel to be configured. However, for vector search, you’ll need an embeddingModel. A name is helpful to attribute each message to a specific agent. Other options are defaults that can be over-ridden at each LLM call-site.