Skip to main content
If you are wiring Synap into a live agent (a chat assistant, a copilot, a voice bot, anything with a conversation loop), this is the integration to build. Your agent reports what happened; Synap handles retrieval and memory formation.
Agent integration flow: sdk.initialize() initializes the SDK, then sdk.instance.listen() starts listening once at startup, then a repeating per-turn block of send_message(user_message), fetch() served from cache, your LLM, and send_message(assistant_message), and finally sdk.instance.stop_listening() on shutdown
That is the whole loop. There is no memories.create() in it.

Why there is no ingestion call

Conversation turns you report with send_message() are persisted to conversation history. When the conversation compacts, Synap promotes those raw turns into the long-term ingestion pipeline: the same extraction stages memories.create() runs, producing memories of the same quality. Compaction happens on any of three triggers: The token and message thresholds are configurable per Instance. The idle trigger is what makes this complete: a two-turn exchange that never approaches the other thresholds still compacts a few minutes after the user stops, and its turns still become memory.
Because the idle trigger catches whatever the thresholds miss, every conversation reaches long-term memory eventually. Reporting turns with send_message() is sufficient on its own.

The integration

1

Initialize once, at process startup

2

Open one stream for the process

Not one per user, not one per session. Scope travels on each call, not on the stream.
This step is the one JavaScript restriction. Every other call in this loop runs on Edge and Workers; listen() is gRPC, so it needs Node plus @grpc/grpc-js and @grpc/proto-loader. Importing the SDK in an Edge route stays safe, only the stream is unavailable there.
See Real-Time Anticipation in a Server for quotas, reconnects, and the health check worth alerting on.
3

Report each turn

Emit assistant_message after the reply. Anticipation runs between turns, so this event is what pre-warms the next one.
4

Close on shutdown

Report the whole turn, not just the two ends

Most integrations report the user’s message and the assistant’s reply and stop there. That works, and it leaves most of the value on the table. Synap then sees the question and the answer and nothing about the work in between, which is where the next question is usually decided. Five things happen in a turn. Report all five.
If you implement only one of these beyond the user turn, implement assistant_message. Anticipation fires on it. An integration that reports tool calls and reasoning but never the reply gets no prefetching at all, and looks from the outside exactly like one that is working.

Where each call goes

No session call appears in that loop. The SDK opens one on the first event and closes it when you stop listening. end_session is only for a conversation that ends while the process lives on: a call hangs up, a chat window closes.

Check your own work

An integration is finished when all of these are true. If you are an agent wiring this up, check them before you report back.
  • listen() is called once at startup, and stop_listening() on shutdown.
  • Every turn reports a user_message and an assistant_message.
  • Tool calls and results are reported, and share a tool_call_id.
  • Reasoning is reported where the framework exposes it.
  • Every event carries conversation_id and user_id (plus customer_id on B2B).
  • No send_message call sets role by hand for a tool or reasoning event. Use the typed methods; the role and the event type have to agree, and getting that wrong fails silently.
  • end_session is called when a conversation ends and the process keeps running.
Using one of the supported frameworks? Do not write any of this by hand. The integration packages for OpenAI Agents, Google ADK, the Claude Agent SDK and the Vercel AI SDK report these events for you from the framework’s own hooks. See Framework integrations.

Requirements

Every event needs both user_id and customer_id. A conversation event missing either is discarded server-side with no client-visible error. The turn is never persisted, so it never becomes memory. On a B2C instance send user_id only: customer_id is not accepted there, and an event without one is persisted normally.
required
The stream must stay open across turns. This integration suits servers, workers, and voice sessions, not per-request serverless functions, which cannot hold a connection.
required
Stream quotas are per Instance and per client. Opening one per user session exhausts them under real concurrency.
required
Reused across every turn of a conversation. send_message() does not validate it, but fetch() and every other call do.

What still uses REST

The stream is not the entire transport, and it is not meant to be. Two things go over HTTP:
  • initialize(): resolves your client and instance from the API key
  • fetch() on an anticipation miss: pushed bundles cover what Synap predicted; anything it did not predict is fetched normally
You do not code either differently. fetch() is one call whether it resolves from the local anticipation cache in about a millisecond or falls through to the network.

When to still call memories.create()

Promotion is automatic but not instant. A turn becomes retrievable when its conversation compacts, which for a quiet conversation means about five minutes. Reach for explicit ingestion in three cases:
use memories.create
A fact the very next turn depends on, or a live demo showing memory forming. Waiting for compaction is not an option.
use memories.create
Promotion ingests conversation turns as conversation content. To set document_type, mode, custom metadata, or to write at customer or client scope, ingest explicitly.
use memories.create
Product docs, support tickets, CRM records, and backfills belong in ingestion, not on the stream.
Do not do both on the same text. If you report a turn with send_message() and ingest that same turn with memories.create(), the content is extracted twice, once immediately and once at compaction, costing double and producing overlapping memories the deduplication stage then reconciles.Pick one owner per piece of content: stream conversation turns, ingest everything else.

Framework integrations

Five packages drive the stream for you, from the framework’s own hooks. With any other framework, or none, use the loop above directly: it is plain SDK calls and composes with anything. Each package is silent unless listen() is running, so adding one changes nothing until you open a stream.

How anticipation works

The stream in depth: event types, cache behavior, and failure modes.

Running it in a server

Quotas, reconnects, and the silent failure to alert on.