Skip to contents

Language: English | 简体中文

This vignette describes the stable public interface for embedding codeagent as the engine behind a host application (a Shiny app, an API service, another package). The host keeps full ownership of its UI, tools, skill content, data, and provider credentials; codeagent provides the agent loop, streaming, context compaction, permission gate, and skill loading.

Everything documented here is Backend Contract v1. See Versioning for the stability promise. A runnable reference lives at system.file("examples/backend_integration_demo.R", package = "codeagent").

The boundary

codeagent provides The host owns
agent loop, multi-turn tool-calling its UI
provider abstraction (any ellmer Chat) its Chat (provider / model / key)
streaming + typed callbacks rendering
context compaction domain tools
central permission gate skill content + data
skill loading session storage (optional)

1. Entry: a harness-only client

Pass your own ellmer::Chat and set register_tools = FALSE so none of codeagent’s coding tools (Bash / Write / Edit / Glob / Grep / git / web) are attached. You get the client state and harness pipeline without registered tools or an installed permission gate.

chat <- ellmer::chat_openai_compatible(
  base_url    = Sys.getenv("MY_BASE_URL"),
  model       = Sys.getenv("MY_MODEL"),
  credentials = function() Sys.getenv("MY_API_KEY")
)

client <- codeagent::codeagent_client(
  chat            = chat,
  register_tools  = FALSE,
  permission_mode = "default",   # default | plan | accept_edits | bypass | ...
  cwd             = getwd()
)

codeagent_client() returns a CodeagentClient with $chat, $settings, and $data_shield (NULL unless enabled). If a Data Shield is enabled on a harness-only client, attach the host tools first and then install the returned shield on the chat. For multi-user Shiny apps, create the client inside the server session (for example via codeagent_app(client_factory = )); never share one mutable client across browser sessions.

2. Driving a turn + the callback contract

Use codeagent_stream() (blocking; pumps its own event loop) or codeagent_stream_async() (returns a promise). Rendering happens entirely through typed callbacks – codeagent does not touch your UI.

result <- codeagent::codeagent_stream(
  client, user_input,
  on_delta        = function(text_chunk)       { ... }, # assistant text
  on_thinking     = function(chunk)            { ... }, # thinking content
  on_tool_request = function(x)                { ... }, # id, name, arguments, intent
  on_tool_result  = function(x)                { ... }, # six fields; see below
  on_error        = function(message, recovered) { ... },
  on_usage        = function(usage)            { ... }, # end-of-turn usage
  on_tick         = function()                 { ... }  # sync only; ~100 ms heartbeat
)
# returns invisibly: list(text, usage, stop_reason, finish_reason)

Callback payloads:

Callback Argument
on_delta text_chunk (character)
on_thinking thinking chunk
on_tool_request list(id, name, arguments, intent) – fires before the gate
on_tool_result list(id, name, display, value, is_error, artifact) in that exact order; artifact is appended after the five original fields
on_error message, recovered
on_usage usage object, including n_tokens, model_limit, warning_state, and cost_last
on_tick no arguments; available on synchronous codeagent_stream() only

The asynchronous function resolves to the same four-field result. Provider finish reasons are normalized into stop_reason; the original mapped provider value is retained as finish_reason.

3. Rich tool results (text / table / image / code / diff / error)

A tool’s value is the portable text the model sees and the final fallback for all UIs. To also expose a rich, UI-neutral artifact, return tool_result(). The result carries three deliberately separate channels:

  • artifact: the primary cross-UI protocol (schema, version, kind, status, icon, title, payload);
  • display: an optional official shinychat adapter; non-shinychat hosts should not parse or depend on it;
  • value: portable text for the model and unsupported/malformed artifacts.
my_tool <- ellmer::tool(
  function(name) {
    df <- summarise_something(name)
    codeagent::tool_result(
      sprintf("%d x %d summary", nrow(df), ncol(df)),
      kind    = "table",
      payload = list(df = df),
      title   = "Summary"
    )
  },
  name = "Summarise",
  description = "Summarise a named dataset.",
  arguments = list(
    name = ellmer::type_string("Dataset name.")
  )
)

The version-1 artifact is available as extra$codeagent$artifact on a result and as artifact on the stream event. Hosts should use tool_result_artifact(event_or_result) rather than reaching into either structure directly. It validates the schema and outer shape and, by default, accepts version 1 only. It returns NULL for unknown or malformed artifacts, allowing a safe fallback through tool_result_value(event_or_result).

kind Typical payload
text list(text = )
table list(df = <data.frame>)
image list(images = list(list(mime = , b64 = )), output = )
code list(text = , lang = , filename = , output = )
diff list(old = , new = , path = )
error list(message = , detail = )

A non-Shiny host renders the artifact itself, for example:

on_tool_result <- function(event) {
  artifact <- codeagent::tool_result_artifact(event)
  if (!is.null(artifact) && identical(artifact$kind, "table")) {
    my_render_table(artifact$payload$df)
  } else {
    render_plain_text(codeagent::tool_result_value(event))
  }
}

See vignette("tool-artifacts") for version negotiation, failure behavior, and the browser trust boundary.

4. Host tools + the permission gate

Register your tools the standard ellmer way, then declare each tool’s capability so the central gate governs it like a native tool.

chat$register_tool(my_tool)
codeagent::register_tool_meta("Summarise", capability = "read") # read|write|exec|net

On a harness-only client (register_tools = FALSE) the gate is not installed automatically. Install it after attaching your tools so approvals route to your ask_fn. The public tools argument is the standalone equivalent of settings$tools:

tool_policy <- list(
  overrides    = list(Summarise = "allow"),       # allow | ask | deny
  capabilities = list(exec = "ask", net = "deny")
)

codeagent::install_permission_gate(
  chat,
  permission_mode = "default",
  tools = tool_policy,
  # Alternatively classify tools here instead of register_tool_meta():
  tool_meta = list(Summarise = "read"),
  ask_fn = function(name, input, id = NULL) {
    host_request_approval(id, name, input) # logical or promise<logical>
  }
)

The optional id is the tool-call ID and matches on_tool_request$id. install_permission_gate() is idempotent per chat: another call refreshes its live mode, approval callback, and policy rather than stacking another gate.

Important: an undeclared tool defaults to capability "read" and is allowed without sensitive-operation gating. If your tool executes code, writes files, or accesses the network, declare it as "exec", "write", or "net" so the gate can ask or deny it. Built-in metadata remains authoritative and cannot be downgraded by host registration.

Policy precedence is per-tool overrides, then per-capability policy, then the active permission mode and rules. The complete policy shape is sets / capabilities / overrides; pass it as tools = to the standalone installer or configure it as settings$tools on a full codeagent client.

5. Skills

Place host skills in a supported project-local skill directory using the <name>/SKILL.md format (for example, .btw/skills/my_skill/SKILL.md). Skill content is yours; codeagent only discovers, loads, and injects it.

skills <- codeagent::list_skills_meta(cwd = getwd())
prompt <- codeagent::load_skill_prompt(
  "my_skill", args = "optional arguments", cwd = getwd()
)
hint <- codeagent::build_skill_hint(cwd = getwd(), max_tokens = 1000L)

6. Provider

codeagent_client(chat = ) accepts any ellmer::Chat (OpenAI-compatible, Databricks, Anthropic, Gemini, Bedrock, Azure, and others). The lower-level stream functions also accept a bare ellmer::Chat; list-based $stream_async duck typing exists for tests, not as the client-construction contract. The host owns and supplies credentials; codeagent uses the supplied Chat rather than owning those credentials.

7. Versioning

Backend Contract v1 is the surface below, with the signatures and behavior documented above. Changes follow semantic versioning; breaking changes bump the major and are announced in NEWS.md. The guard test test-backend-contract.R fails if the exported surface drifts.