From 14d59fc925e13b794204cbcccc1f27eb7b704448 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Tue, 18 Nov 2025 14:16:46 -0800 Subject: [PATCH 01/30] Enhance README with tool use modes and examples Added detailed explanations for tool use modes, including examples for open loop and closed loop execution. --- README.md | 74 ++++++++++++++++++++++++++++++++++++++++++++++++++++--- 1 file changed, 71 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index aace534..98b0317 100644 --- a/README.md +++ b/README.md @@ -141,7 +141,74 @@ Because of their special behavior of being preserved on context window overflow, The Prompt API supports **tool use** via the `tools` option, allowing you to define external capabilities that a language model can invoke in a model-agnostic way. Each tool is represented by an object that includes an `execute` member that specifies the JavaScript function to be called. When the language model initiates a tool use request, the user agent calls the corresponding `execute` function and sends the result back to the model. -Here’s an example of how to use the `tools` option: +There are 2 tool use modes: with automatic execution (closed loop) and without automatic execution (open loop) + +Regardless of with or without automatic execution, the session creation and appending signature are the same. Here’s an example: + +```js +const session = await LanguageModel.create({ + initialPrompts: [ + { + role: "system", + content: `You are a helpful assistant. You can use tools to help the user.` + } + ], + tools: [ + { + name: "getWeather", + description: "Get the weather in a location.", + inputSchema: { + type: "object", + properties: { + location: { + type: "string", + description: "The city to check for the weather condition.", + }, + }, + required: ["location"], + }, + } + ] +}); +``` + +In this example, the `tools` array defines a `getWeather` tool, specifying its name, description and input schema. + +Few shot examples of tool use can be appended like so: + +```js +await session.append([ + {role: "user", content: "What is the weather in Seattle?"}, + {role: "tool-call", content: {type: "tool-call", value: {callID:" get_weather_1", name: "get_weather", arguments: {location:"Seattle"}}}, + {role: "tool-result", content: {type: "tool-response", value: {callID: "get_weather_1", name: "get_weather", result: [{type:"object", value: {temperature: "55F", humidity: "67%"}}]}}, + {role: "assistant", content: "The temperature in Seattle is 55F and humidity is 67%"}, +]); +``` + +Note that "role" and "type" now supports "tool-call" and "tool-result". `content.result` is a list of a dictionary of `type` and `value`, where `type` can be `{"text", "image", "audio", "object" }` and `value` is `any`. + +#### Open Loop: + +Without automatic execution, the API will return a `ToolCall` object with `callId` (a unique identifier of this tool call), `name` (name of the tool), and `arguments` (a dictionary fitting the JSON input schema of the tool's declaration), and client is expected to handle the tool execution and append the tool result back to the session. + +Example: + +```js +const result = await session.prompt("What is the weather in Seattle?"); +if (result.type=="tool-call") { + if (result.name == "get_weather") { + const tool_result = getWeather(result.arguments.location); + session.prompt([{role:"tool-result", content: {type: "tool-result", value: {callId: result.callID, name: result.name, result: [{type:"object", value: tool_result}]}}}]) + } +} +``` + +Note that we always require tool-response to immediately follow tool-call generated by the model. + + +#### Closed Loop: + +To enable automatic execution, add a `execute` function for each tool's implementation, and add a `toolUseConfig` to indicate that execution is enabled and pose a max number of tool calls invoked in a single session generation: ```js const session = await LanguageModel.create({ @@ -171,13 +238,14 @@ const session = await LanguageModel.create({ return JSON.stringify(await res.json()); }, } - ] + ], +toolUseConfig: {enabled: true, max_tool_calls: 5}, }); const result = await session.prompt("What is the weather in Seattle?"); ``` -In this example, the `tools` array defines a `getWeather` tool, specifying its name, description, input schema, and `execute` implementation. When the language model determines that a tool call is needed, the user agent invokes the `getWeather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. +When the language model determines that a tool call is needed, the user agent invokes the `getWeather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. #### Concurrent tool use From 5379da1fc97469d89867530ef2fb144a31fbdda4 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Wed, 19 Nov 2025 09:53:15 -0800 Subject: [PATCH 02/30] Clarify tool-call and tool-result in README Updated README to clarify tool-call and tool-result usage. --- README.md | 19 ++++++++++++++----- 1 file changed, 14 insertions(+), 5 deletions(-) diff --git a/README.md b/README.md index 98b0317..6bcf806 100644 --- a/README.md +++ b/README.md @@ -185,21 +185,30 @@ await session.append([ ]); ``` -Note that "role" and "type" now supports "tool-call" and "tool-result". `content.result` is a list of a dictionary of `type` and `value`, where `type` can be `{"text", "image", "audio", "object" }` and `value` is `any`. +Note that "role" and "type" now supports "tool-call" and "tool-result". +`content.result` is a list of a dictionary of `type` and `value`, where `type` can be `{"text", "image", "audio", "object" }` and `value` is `any`. #### Open Loop: -Without automatic execution, the API will return a `ToolCall` object with `callId` (a unique identifier of this tool call), `name` (name of the tool), and `arguments` (a dictionary fitting the JSON input schema of the tool's declaration), and client is expected to handle the tool execution and append the tool result back to the session. +Open loop is enabled by specifying `tool-call` in `expectedOutputs` when the session is created. + +When a tool needs to be called, the API will return an object with `callId` (a unique identifier of this tool call), `name` (name of the tool), and `arguments` (inputs to the tool), and client is expected to handle the tool execution and append the tool result back to the session. The `argument` is a dictionary fitting the JSON input schema of the tool's declaration; if the input schema is not "object", the value will be wrapped in a key. Example: ```js -const result = await session.prompt("What is the weather in Seattle?"); +sessionOptions = structuredClone(options); +sessionOptions.expectedOutputs.push(["tool-call"]); +session = await LanguageModel.create(sessionOptions); + +var result = await session.prompt("What is the weather in Seattle?"); if (result.type=="tool-call") { if (result.name == "get_weather") { const tool_result = getWeather(result.arguments.location); - session.prompt([{role:"tool-result", content: {type: "tool-result", value: {callId: result.callID, name: result.name, result: [{type:"object", value: tool_result}]}}}]) + result = session.prompt([{role:"tool-result", content: {type: "tool-result", value: {callId: result.callID, name: result.name, result: [{type:"object", value: tool_result}]}}}]) } +} else{ + console.log(result) } ``` @@ -239,7 +248,7 @@ const session = await LanguageModel.create({ }, } ], -toolUseConfig: {enabled: true, max_tool_calls: 5}, +toolUseConfig: {enabled: true}, }); const result = await session.prompt("What is the weather in Seattle?"); From 8e076f7f3bcad312527d7bbbf458a0f4aa5f700d Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Wed, 19 Nov 2025 10:00:49 -0800 Subject: [PATCH 03/30] Enhance LanguageModel with tool call definitions Added new types and enums for tool calls and responses. --- index.bs | 51 ++++++++++++++++++++++++++++++++++++++++++++------- 1 file changed, 44 insertions(+), 7 deletions(-) diff --git a/index.bs b/index.bs index 7f7bdb1..37c8ebc 100644 --- a/index.bs +++ b/index.bs @@ -39,8 +39,11 @@ interface LanguageModel : EventTarget { static Promise availability(optional LanguageModelCreateCoreOptions options = {}); static Promise params(); + // The return type from prompt() method and those alike. + typedef (DOMString or sequence) LanguageModelPromptResult; + // These will throw "NotSupportedError" DOMExceptions if role = "system" - Promise prompt( + Promise prompt( LanguageModelPrompt input, optional LanguageModelPromptOptions options = {} ); @@ -80,13 +83,11 @@ interface LanguageModelParams { callback LanguageModelToolFunction = Promise (any... arguments); // A description of a tool call that a language model can invoke. -dictionary LanguageModelTool { +dictionary LanguageModelToolDeclaration { required DOMString name; required DOMString description; // JSON schema for the input parameters. required object inputSchema; - // The function to be invoked by user agent on behalf of language model. - required LanguageModelToolFunction execute; }; dictionary LanguageModelCreateCoreOptions { @@ -97,7 +98,7 @@ dictionary LanguageModelCreateCoreOptions { sequence expectedInputs; sequence expectedOutputs; - sequence tools; + sequence tools; }; dictionary LanguageModelCreateOptions : LanguageModelCreateCoreOptions { @@ -148,16 +149,52 @@ dictionary LanguageModelMessageContent { required LanguageModelMessageValue value; }; -enum LanguageModelMessageRole { "system", "user", "assistant" }; +enum LanguageModelMessageRole { "system", "user", "assistant", "tool-call", "tool-response" }; -enum LanguageModelMessageType { "text", "image", "audio" }; +enum LanguageModelMessageType { "text", "image", "audio","tool-call", "tool-response" }; typedef ( ImageBitmapSource or AudioBuffer or BufferSource or DOMString + or LanguageModelToolCall + or LanguageModelToolResponse ) LanguageModelMessageValue; + +// The definitions of `LanguageModelToolCall` and `LanguageModelToolResponse` values +enum LanguageModelToolResultType { "text", "image", "audio", "object" }; + +dictionary LanguageModelToolResultContent { + required LanguageModelToolResultType type; + required any value; +}; + +// Represents a tool call requested by the language model. +dictionary LanguageModelToolCall { + required DOMString callID; + required DOMString name; + object arguments; +}; + +// Successful tool execution result. +dictionary LanguageModelToolSuccess { + required DOMString callID; + required DOMString name; + required sequence result; +}; + +// Failed tool execution result. +dictionary LanguageModelToolError { + required DOMString callID; + required DOMString name; + required DOMString errorMessage; +}; + +// The response from executing a tool call - either success or error. +typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolResponse; + +

Prompt processing

From ed78c121f50a45af2f8ca9cd25613cb68677032b Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Wed, 26 Nov 2025 09:08:32 -0800 Subject: [PATCH 04/30] Update README.md Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 6bcf806..d4aaea4 100644 --- a/README.md +++ b/README.md @@ -141,7 +141,7 @@ Because of their special behavior of being preserved on context window overflow, The Prompt API supports **tool use** via the `tools` option, allowing you to define external capabilities that a language model can invoke in a model-agnostic way. Each tool is represented by an object that includes an `execute` member that specifies the JavaScript function to be called. When the language model initiates a tool use request, the user agent calls the corresponding `execute` function and sends the result back to the model. -There are 2 tool use modes: with automatic execution (closed loop) and without automatic execution (open loop) +There are two tool use modes: with automatic execution (closed loop) and without automatic execution (open loop). Regardless of with or without automatic execution, the session creation and appending signature are the same. Here’s an example: From 5b11a1beba282caac1cad95973a676e15c8762e2 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Wed, 26 Nov 2025 09:08:43 -0800 Subject: [PATCH 05/30] Update README.md Co-authored-by: Thomas Steiner --- README.md | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index d4aaea4..804f4bd 100644 --- a/README.md +++ b/README.md @@ -150,8 +150,8 @@ const session = await LanguageModel.create({ initialPrompts: [ { role: "system", - content: `You are a helpful assistant. You can use tools to help the user.` - } + content: `You are a helpful assistant. You can use tools to help the user.`, + }, ], tools: [ { @@ -167,8 +167,8 @@ const session = await LanguageModel.create({ }, required: ["location"], }, - } - ] + }, + ], }); ``` From 5a0cbefcc650dd756f38ae26b546f9deaf16c0a4 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Wed, 26 Nov 2025 11:49:37 -0800 Subject: [PATCH 06/30] add open vs. loop use case diff --- README.md | 24 ++++++++++++++++++++++++ 1 file changed, 24 insertions(+) diff --git a/README.md b/README.md index 804f4bd..18f8788 100644 --- a/README.md +++ b/README.md @@ -256,6 +256,30 @@ const result = await session.prompt("What is the weather in Seattle?"); When the language model determines that a tool call is needed, the user agent invokes the `getWeather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. +#### Do I need auto execution? + +In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you are tolerable for certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g, just a few distinct tools, short and clean tool output, short context window, etc) + +On the other hand, open loop allows more flexibility for intercepting at various points in the planner loop (the reason->action->observation loop) where you can inject your business logic programmatically. + +Here are a few patterns where open loop would be useful: + +1) context management + +If your session might go through a long chain of contents, and the previous tool results are no longer important or relevant for your use case, open loop gives the flexibility of editing and recreating the session in the middle of a tool call. You can manually compress and modify the history, and recreate a new session with less content. + +For example, for a shopping agent, your tool keeps track of a live shopping cart, but only the latest cart status is important. When there have been multiple rounds of cart updates, you might need to compress the tool call history to avoid exceeding context window, improve latency and quality. + +2) Conditional loop breaking + +If your business logic requires some determinism in some critical states, open loop allows the flexibility to early exit the planner loop and output a pre-determined action. + +For example, for a shopping agent, you might be required to get an explicit confirmation before placing the order. Whenever the tool `"place_order"` is called in the first time, you want to exit the planner loop immediately, and display a verbatim message to the user + +3) Conditional constraints + + + #### Concurrent tool use Developers should be aware that the model might call their tool multiple times, concurrently. For example, code such as From 81e75afb93de8a640b5791f8027b5941991f4d7a Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 1 Dec 2025 09:25:32 -0800 Subject: [PATCH 07/30] Enhance README with planner loop constraints details Added explanation about automatic execution and constraints in planner loop. --- README.md | 5 ++++- 1 file changed, 4 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index 18f8788..814f333 100644 --- a/README.md +++ b/README.md @@ -277,7 +277,10 @@ If your business logic requires some determinism in some critical states, open l For example, for a shopping agent, you might be required to get an explicit confirmation before placing the order. Whenever the tool `"place_order"` is called in the first time, you want to exit the planner loop immediately, and display a verbatim message to the user 3) Conditional constraints - + +In automatic execution, the planner loop decodes various and mutliple times. If you need to supply constraints dynamically, you'd use the open loop API and control the planner loop yourself. Because the planner loop runs the entire loop behind the scene, the closed loop API doesn't have a natural way to supply a different constraint for each LLM step. + +For example, you might want the model to always generate tool `FOO` after tool `BAR` is called; or you might want the model to always generate text only with some prefix after tool `FOO` is called. #### Concurrent tool use From 7afff121264138ef5afb447354afb7c45c217f44 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:33:45 -0700 Subject: [PATCH 08/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 33 +++++++++++++++++++++++++++++---- 1 file changed, 29 insertions(+), 4 deletions(-) diff --git a/README.md b/README.md index 71f4507..29bdffa 100644 --- a/README.md +++ b/README.md @@ -212,10 +212,35 @@ Few shot examples of tool use can be appended like so: ```js await session.append([ - {role: "user", content: "What is the weather in Seattle?"}, - {role: "tool-call", content: {type: "tool-call", value: {callID:" get_weather_1", name: "get_weather", arguments: {location:"Seattle"}}}, - {role: "tool-result", content: {type: "tool-response", value: {callID: "get_weather_1", name: "get_weather", result: [{type:"object", value: {temperature: "55F", humidity: "67%"}}]}}, - {role: "assistant", content: "The temperature in Seattle is 55F and humidity is 67%"}, + { role: "user", content: "What is the weather in Seattle?" }, + { + role: "tool-call", + content: { + type: "tool-call", + value: { + callID: " get_weather_1", + name: "get_weather", + arguments: { location: "Seattle" }, + }, + }, + }, + { + role: "tool-result", + content: { + type: "tool-response", + value: { + callID: "get_weather_1", + name: "get_weather", + result: [ + { type: "object", value: { temperature: "55F", humidity: "67%" } }, + ], + }, + }, + }, + { + role: "assistant", + content: "The temperature in Seattle is 55F and humidity is 67%", + }, ]); ``` From fd46bb60b64178483cb9d21941f09f73a21d993d Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:34:01 -0700 Subject: [PATCH 09/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 29bdffa..9a6b12d 100644 --- a/README.md +++ b/README.md @@ -256,7 +256,7 @@ When a tool needs to be called, the API will return an object with `callId` (a u Example: ```js -sessionOptions = structuredClone(options); +const sessionOptions = structuredClone(options); sessionOptions.expectedOutputs.push(["tool-call"]); session = await LanguageModel.create(sessionOptions); From 8ddbb238a2e227c4f12e7b298506dbb5ea6d474c Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:34:09 -0700 Subject: [PATCH 10/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 9a6b12d..d4e9fbe 100644 --- a/README.md +++ b/README.md @@ -260,7 +260,7 @@ const sessionOptions = structuredClone(options); sessionOptions.expectedOutputs.push(["tool-call"]); session = await LanguageModel.create(sessionOptions); -var result = await session.prompt("What is the weather in Seattle?"); +let result = await session.prompt("What is the weather in Seattle?"); if (result.type=="tool-call") { if (result.name == "get_weather") { const tool_result = getWeather(result.arguments.location); From b9bcf7b2dd200e2b310b0a1131c25fe3b4a27dd4 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:34:17 -0700 Subject: [PATCH 11/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index d4e9fbe..3c46a06 100644 --- a/README.md +++ b/README.md @@ -258,7 +258,7 @@ Example: ```js const sessionOptions = structuredClone(options); sessionOptions.expectedOutputs.push(["tool-call"]); -session = await LanguageModel.create(sessionOptions); +const session = await LanguageModel.create(sessionOptions); let result = await session.prompt("What is the weather in Seattle?"); if (result.type=="tool-call") { From b04dec1ebf75720b128a41906d09b97e69bdd05e Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:34:24 -0700 Subject: [PATCH 12/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 3c46a06..5cbba1b 100644 --- a/README.md +++ b/README.md @@ -276,7 +276,7 @@ Note that we always require tool-response to immediately follow tool-call genera #### Closed Loop: -To enable automatic execution, add a `execute` function for each tool's implementation, and add a `toolUseConfig` to indicate that execution is enabled and pose a max number of tool calls invoked in a single session generation: +To enable automatic execution, add an `execute` function for each tool's implementation, and add a `toolUseConfig` to indicate that execution is enabled and pose a max number of tool calls invoked in a single session generation: ```js const session = await LanguageModel.create({ From 0e6f916ff5903a8ca8ee4ac97a44a794e537e6b5 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 15:34:47 -0700 Subject: [PATCH 13/30] Apply suggestion from @tomayac Co-authored-by: Thomas Steiner --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 5cbba1b..716d2cc 100644 --- a/README.md +++ b/README.md @@ -244,7 +244,7 @@ await session.append([ ]); ``` -Note that "role" and "type" now supports "tool-call" and "tool-result". +Note that `"role"` and `"type"` now support `"tool-call"` and `"tool-result"`. `content.result` is a list of a dictionary of `type` and `value`, where `type` can be `{"text", "image", "audio", "object" }` and `value` is `any`. #### Open Loop: From 3e35c8908f5e7909698ec1f8c0eebfb20f155baa Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 16:10:49 -0700 Subject: [PATCH 14/30] Update tool use documentation in explainer to match implementation --- README.md | 159 ++++++++++++++++++++++++++++++++---------------------- 1 file changed, 95 insertions(+), 64 deletions(-) diff --git a/README.md b/README.md index 716d2cc..e604041 100644 --- a/README.md +++ b/README.md @@ -173,23 +173,25 @@ Because of their special behavior of being preserved on context window overflow, ### Tool use -The Prompt API supports **tool use** via the `tools` option, allowing you to define external capabilities that a language model can invoke in a model-agnostic way. Each tool is represented by an object that includes an `execute` member that specifies the JavaScript function to be called. When the language model initiates a tool use request, the user agent calls the corresponding `execute` function and sends the result back to the model. +The Prompt API supports **tool use** via the `tools` option, allowing you to define external capabilities that a language model can invoke in a model-agnostic way. Each tool is declared with a `name`, `description`, and `inputSchema` (a JSON Schema object with `type: "object"`). -There are two tool use modes: with automatic execution (closed loop) and without automatic execution (open loop). +There are two tool use modes: without automatic execution (**open loop**) and with automatic execution (**closed loop**, planned). In open loop mode, the model returns tool call requests to the application, which executes the tool and sends the response back to the session. In closed loop mode, each tool declaration also includes an `execute` function that the user agent invokes automatically when the model requests a tool call. -Regardless of with or without automatic execution, the session creation and appending signature are the same. Here’s an example: +Here is an example of creating a session with tool declarations: ```js const session = await LanguageModel.create({ initialPrompts: [ { role: "system", - content: `You are a helpful assistant. You can use tools to help the user.`, + content: "You are a helpful assistant. You can use tools to help the user.", }, ], + expectedInputs: [{ type: "tool-call" }, { type: "tool-response" }], + expectedOutputs: [{ type: "tool-call" }], tools: [ { - name: "getWeather", + name: "get_weather", description: "Get the weather in a location.", inputSchema: { type: "object", @@ -206,75 +208,102 @@ const session = await LanguageModel.create({ }); ``` -In this example, the `tools` array defines a `getWeather` tool, specifying its name, description and input schema. +In this example, the `tools` array defines a `get_weather` tool, specifying its name, description, and input schema. When `tools` are provided, `expectedOutputs` must include `{ type: "tool-call" }`. To supply tool calls and/or tool responses in `initialPrompts`, `append()`, or `prompt()`, `expectedInputs` must also include `{ type: "tool-call" }` and/or `{ type: "tool-response" }`. -Few shot examples of tool use can be appended like so: +Few-shot examples of tool use can be appended like so: ```js await session.append([ { role: "user", content: "What is the weather in Seattle?" }, { - role: "tool-call", - content: { - type: "tool-call", - value: { - callID: " get_weather_1", - name: "get_weather", - arguments: { location: "Seattle" }, + role: "assistant", + content: [ + { + type: "tool-call", + value: new LanguageModelToolCall({ + callID: "get_weather_1", + name: "get_weather", + arguments: { location: "Seattle" }, + }), }, - }, + ], }, { - role: "tool-result", - content: { - type: "tool-response", - value: { - callID: "get_weather_1", - name: "get_weather", - result: [ - { type: "object", value: { temperature: "55F", humidity: "67%" } }, - ], + role: "user", + content: [ + { + type: "tool-response", + value: new LanguageModelToolSuccess({ + callID: "get_weather_1", + name: "get_weather", + result: [ + { type: "object", value: { temperature: "55F", humidity: "67%" } }, + ], + }), }, - }, + ], }, { role: "assistant", - content: "The temperature in Seattle is 55F and humidity is 67%", + content: "The temperature in Seattle is 55F and humidity is 67%.", }, ]); ``` -Note that `"role"` and `"type"` now support `"tool-call"` and `"tool-result"`. -`content.result` is a list of a dictionary of `type` and `value`, where `type` can be `{"text", "image", "audio", "object" }` and `value` is `any`. +Note that: +* Message `content` `type` supports `"tool-call"` and `"tool-response"`: + * `"tool-call"` content must use `role: "assistant"` and its `value` must be a `LanguageModelToolCall` instance (`new LanguageModelToolCall({ callID, name, arguments })`). + * `"tool-response"` content must use `role: "user"` and its `value` must be either a `LanguageModelToolSuccess` instance (`new LanguageModelToolSuccess({ callID, name, result })`) or a `LanguageModelToolError` instance (`new LanguageModelToolError({ callID, name, errorMessage })`). +* `LanguageModelToolSuccess.result` is a list of `{ type, value }` dictionaries, where `type` can be `"text"`, `"image"`, `"audio"`, or `"object"`, and `value` is `any`. -#### Open Loop: +#### Open Loop -Open loop is enabled by specifying `tool-call` in `expectedOutputs` when the session is created. +Open loop is enabled by specifying `{ type: "tool-call" }` in `expectedOutputs` when the session is created. -When a tool needs to be called, the API will return an object with `callId` (a unique identifier of this tool call), `name` (name of the tool), and `arguments` (inputs to the tool), and client is expected to handle the tool execution and append the tool result back to the session. The `argument` is a dictionary fitting the JSON input schema of the tool's declaration; if the input schema is not "object", the value will be wrapped in a key. +When the model does not invoke any tools, `session.prompt()` resolves to a `DOMString` as usual. When a tool needs to be called, `session.prompt()` resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence`). If the model outputs both text and tool calls, it's resolved to an array, where the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. -Example: +Each `LanguageModelToolCall` object contains: +* `callID`: A unique string identifier for this tool call. +* `name`: The name of the tool to invoke. +* `arguments`: A dictionary fitting the JSON `inputSchema` of the tool's declaration (which must have `type: "object"`). + +The client is expected to execute the tool and send the result back to the session using `LanguageModelToolSuccess` (or `LanguageModelToolError` if execution failed): ```js const sessionOptions = structuredClone(options); -sessionOptions.expectedOutputs.push(["tool-call"]); +sessionOptions.expectedInputs = [{ type: "tool-response" }]; +sessionOptions.expectedOutputs = [{ type: "tool-call" }]; const session = await LanguageModel.create(sessionOptions); let result = await session.prompt("What is the weather in Seattle?"); -if (result.type=="tool-call") { - if (result.name == "get_weather") { - const tool_result = getWeather(result.arguments.location); - result = session.prompt([{role:"tool-result", content: {type: "tool-result", value: {callId: result.callID, name: result.name, result: [{type:"object", value: tool_result}]}}}]) +if (Array.isArray(result)) { + const toolCallMsg = result.find((msg) => msg.type === "tool-call"); + if (toolCallMsg && toolCallMsg.value.name === "get_weather") { + const toolCall = toolCallMsg.value; + const toolResult = getWeather(toolCall.arguments.location); + result = await session.prompt([ + { + role: "user", + content: [ + { + type: "tool-response", + value: new LanguageModelToolSuccess({ + callID: toolCall.callID, + name: toolCall.name, + result: [{ type: "object", value: toolResult }], + }), + }, + ], + }, + ]); } -} else{ - console.log(result) } +console.log(result); ``` -Note that we always require tool-response to immediately follow tool-call generated by the model. +Note that a `"tool-response"` message should immediately follow the `"tool-call"` generated by the model. - -#### Closed Loop: +#### Closed Loop To enable automatic execution, add an `execute` function for each tool's implementation, and add a `toolUseConfig` to indicate that execution is enabled and pose a max number of tool calls invoked in a single session generation: @@ -283,14 +312,14 @@ const session = await LanguageModel.create({ initialPrompts: [ { role: "system", - content: `You are a helpful assistant. You can use tools to help the user.`, + content: "You are a helpful assistant. You can use tools to help the user.", }, ], expectedInputs: [{ type: "text", languages: ["en"] }, { type: "tool-response" }], expectedOutputs: [{ type: "text", languages: ["en"] }, { type: "tool-call" }], tools: [ { - name: "getWeather", + name: "get_weather", description: "Get the weather in a location.", inputSchema: { type: "object", @@ -304,7 +333,7 @@ const session = await LanguageModel.create({ }, async execute({ location }) { const res = await fetch( - "https://weatherapi.example/?location=" + location, + "https://weatherapi.example/?location=" + encodeURIComponent(location), ); // Returns the result as a JSON string. return JSON.stringify(await res.json()); @@ -316,46 +345,48 @@ const session = await LanguageModel.create({ const result = await session.prompt("What is the weather in Seattle?"); ``` -When the language model determines that a tool call is needed, the user agent invokes the `getWeather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. +When the language model determines that a tool call is needed, the user agent invokes the `get_weather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. #### Do I need auto execution? -In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you are tolerable for certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g, just a few distinct tools, short and clean tool output, short context window, etc) +In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you can tolerate certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g., just a few distinct tools, short and clean tool output, short context window, etc.). + +On the other hand, open loop allows more flexibility for intercepting at various points in the planner loop (the reason -> action -> observation loop) where you can inject your business logic programmatically. -On the other hand, open loop allows more flexibility for intercepting at various points in the planner loop (the reason->action->observation loop) where you can inject your business logic programmatically. +Here are a few patterns where open loop is useful: -Here are a few patterns where open loop would be useful: +1) **Context management** -1) context management +If your session might go through a long chain of interactions, and the previous tool results are no longer important or relevant for your use case, open loop gives you the flexibility of editing and recreating the session in the middle of a tool call. You can manually compress and modify the history, and recreate a new session with less content. -If your session might go through a long chain of contents, and the previous tool results are no longer important or relevant for your use case, open loop gives the flexibility of editing and recreating the session in the middle of a tool call. You can manually compress and modify the history, and recreate a new session with less content. +For example, for a shopping agent, your tool keeps track of a live shopping cart, but only the latest cart status is important. When there have been multiple rounds of cart updates, you might need to compress the tool call history to avoid exceeding the context window and to improve latency and quality. -For example, for a shopping agent, your tool keeps track of a live shopping cart, but only the latest cart status is important. When there have been multiple rounds of cart updates, you might need to compress the tool call history to avoid exceeding context window, improve latency and quality. - -2) Conditional loop breaking +2) **Conditional loop breaking** -If your business logic requires some determinism in some critical states, open loop allows the flexibility to early exit the planner loop and output a pre-determined action. +If your business logic requires determinism in some critical states, open loop allows the flexibility to early-exit the planner loop and output a pre-determined action. -For example, for a shopping agent, you might be required to get an explicit confirmation before placing the order. Whenever the tool `"place_order"` is called in the first time, you want to exit the planner loop immediately, and display a verbatim message to the user - -3) Conditional constraints +For example, for a shopping agent, you might be required to get an explicit confirmation before placing the order. Whenever the tool `"place_order"` is called for the first time, you can exit the planner loop immediately and display a verbatim confirmation message to the user. -In automatic execution, the planner loop decodes various and mutliple times. If you need to supply constraints dynamically, you'd use the open loop API and control the planner loop yourself. Because the planner loop runs the entire loop behind the scene, the closed loop API doesn't have a natural way to supply a different constraint for each LLM step. +3) **Conditional constraints** -For example, you might want the model to always generate tool `FOO` after tool `BAR` is called; or you might want the model to always generate text only with some prefix after tool `FOO` is called. +In automatic execution, the planner loop decodes multiple times behind the scenes. If you need to supply constraints dynamically, you can use the open loop API and control the planner loop yourself, since closed loop doesn't have a natural way to supply a different `responseConstraint` for each LLM step. +For example, you might want the model to always generate tool `FOO` after tool `BAR` is called, or you might want the model to always generate text-only output with a specific prefix after tool `FOO` is called. #### Concurrent tool use -Developers should be aware that the model might call their tool multiple times, concurrently. For example, code such as +Developers should be aware that the model might call their tool multiple times, concurrently. For example, code such as: ```js -const result = await session.prompt("Which of these locations currently has the highest temperature? Seattle, Tokyo, Berlin"); +const result = await session.prompt( + "Which of these locations currently has the highest temperature? Seattle, Tokyo, Berlin" +); ``` -might call the above `"getWeather"` tool's `execute()` function three times. The model would wait for all tool call results to return, using the equivalent of `Promise.all()` internally, before it composes its final response. +might emit three `"get_weather"` tool calls in a single turn (or, in closed loop mode, call `execute()` three times concurrently and wait for all tool call results using the equivalent of `Promise.all()` internally before composing its final response). + +Similarly, the model might call multiple different tools in a single turn if it believes they are all relevant when responding to the given prompt. -Similarly, the model might call multiple different tools, if it believes they all are relevant when responding to the given prompt. ### Multimodal inputs From a5b01a37f6e618d75d2d3d32a393a95ad14f0a55 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Mon, 28 Sep 2026 16:23:21 -0700 Subject: [PATCH 15/30] Update tool use spec based on IDL and implementation --- index.bs | 44 ++++++++++++++++++++++++++++++-------------- 1 file changed, 30 insertions(+), 14 deletions(-) diff --git a/index.bs b/index.bs index a49be78..90b9f22 100644 --- a/index.bs +++ b/index.bs @@ -42,6 +42,9 @@ These APIs are part of a family of APIs expected to be powered by machine learni

The API

+// The return type from prompt() method and those alike. +typedef (DOMString or sequence<LanguageModelMessageContent>) LanguageModelPromptResult; + [Exposed=Window, SecureContext] interface LanguageModel : EventTarget { static Promise<LanguageModel> create(optional LanguageModelCreateOptions options = {}); @@ -49,9 +52,6 @@ interface LanguageModel : EventTarget { // **EXPERIMENTAL**: Only available in extension and experimental contexts. static Promise<LanguageModelParams?> params(); - // The return type from prompt() method and those alike. - typedef (DOMString or sequence<LanguageModelMessageContent>) LanguageModelPromptResult; - // These will throw "NotSupportedError" DOMExceptions if role = "system" Promise<LanguageModelPromptResult> prompt( LanguageModelPrompt input, @@ -107,8 +107,6 @@ interface LanguageModelParams { readonly attribute float maxTemperature; }; -callback LanguageModelToolFunction = Promise<DOMString> (any... arguments); - // A description of a tool call that a language model can invoke. dictionary LanguageModelToolDeclaration { required DOMString name; @@ -188,7 +186,7 @@ dictionary LanguageModelMessageContent { enum LanguageModelSamplingMode { "most-predictable", "predictable", "slightly-predictable", "balanced", "slightly-creative", "creative", "most-creative" }; -enum LanguageModelMessageRole { "system", "user", "assistant", "tool-call", "tool-response" }; +enum LanguageModelMessageRole { "system", "user", "assistant" }; enum LanguageModelMessageType { "text", "image", "audio", "tool-call", "tool-response" }; @@ -210,21 +208,39 @@ dictionary LanguageModelToolResultContent { }; // Represents a tool call requested by the language model. -dictionary LanguageModelToolCall { +[Exposed=Window, SecureContext] +interface LanguageModelToolCall { + constructor(LanguageModelToolCallInit init); + readonly attribute DOMString callID; + readonly attribute DOMString name; + readonly attribute object? arguments; +}; + +dictionary LanguageModelToolCallInit { required DOMString callID; required DOMString name; object arguments; }; - -// Successful tool execution result. -dictionary LanguageModelToolSuccess { +[Exposed=Window, SecureContext] +interface LanguageModelToolSuccess { + constructor(LanguageModelToolSuccessInit init); + readonly attribute DOMString callID; + readonly attribute DOMString name; + readonly attribute FrozenArray<LanguageModelToolResultContent> result; +}; +dictionary LanguageModelToolSuccessInit { required DOMString callID; required DOMString name; required sequence<LanguageModelToolResultContent> result; }; - -// Failed tool execution result. -dictionary LanguageModelToolError { +[Exposed=Window, SecureContext] +interface LanguageModelToolError { + constructor(LanguageModelToolErrorInit init); + readonly attribute DOMString callID; + readonly attribute DOMString name; + readonly attribute DOMString errorMessage; +}; +dictionary LanguageModelToolErrorInit { required DOMString callID; required DOMString name; required DOMString errorMessage; @@ -442,7 +458,7 @@ Every {{LanguageModel}} has an <dfn for="LanguageModel">expected inputs</dfn>, a Every {{LanguageModel}} has an <dfn for="LanguageModel">expected outputs</dfn>, a [=list=] of {{LanguageModelExpected}}s, set during creation. -Every {{LanguageModel}} has a <dfn for="LanguageModel">tools</dfn>, a [=list=] of {{LanguageModelTool}}s, set during creation. +Every {{LanguageModel}} has a <dfn for="LanguageModel">tools</dfn>, a [=list=] of {{LanguageModelToolDeclaration}}s, set during creation. Every {{LanguageModel}} has a <dfn for="LanguageModel">context window size</dfn>, an unrestricted double, set during creation. From ef39f5d05f2a9543046de807a66f1dadf67888e9 Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Tue, 29 Sep 2026 11:46:38 -0700 Subject: [PATCH 16/30] update algorithms for tool call in spec --- index.bs | 190 ++++++++++++++++++++++++++++++++++++++++++++++++------- 1 file changed, 167 insertions(+), 23 deletions(-) diff --git a/index.bs b/index.bs index 90b9f22..e289ba6 100644 --- a/index.bs +++ b/index.bs @@ -52,7 +52,7 @@ interface LanguageModel : EventTarget { // **EXPERIMENTAL**: Only available in extension and experimental contexts. static Promise<LanguageModelParams?> params(); - // These will throw "NotSupportedError" DOMExceptions if role = "system" + // These will throw "TypeError" DOMExceptions if role = "system" Promise<LanguageModelPromptResult> prompt( LanguageModelPrompt input, optional LanguageModelPromptOptions options = {} @@ -266,10 +266,30 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. If |options|["{{LanguageModelCreateCoreOptions/samplingMode}}"] [=map/exists=] and either |options|["{{LanguageModelCreateCoreOptions/topK}}"] [=map/exists=] or |options|["{{LanguageModelCreateCoreOptions/temperature}}"] [=map/exists=], then throw a {{TypeError}}. 1. If |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] [=map/exists=], then [=list/for each=] |expected| of |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"]: - 1. If |expected|["{{LanguageModelExpected/languages}}"] [=map/exists=], then [=Validate and canonicalize language tags=] given |expected| and "{{LanguageModelExpected/languages}}". + 1. If |expected|["{{LanguageModelExpected/languages}}"] [=map/exists=], then [=validate and canonicalize language tags=] given |expected| and "{{LanguageModelExpected/languages}}". + + 1. Let |hasToolCallInExpectedOutputs| be false. 1. If |options|["{{LanguageModelCreateCoreOptions/expectedOutputs}}"] [=map/exists=], then [=list/for each=] |expected| of |options|["{{LanguageModelCreateCoreOptions/expectedOutputs}}"]: - 1. If |expected|["{{LanguageModelExpected/languages}}"] [=map/exists=], then [=Validate and canonicalize language tags=] given |expected| and "{{LanguageModelExpected/languages}}". + 1. If |expected|["{{LanguageModelExpected/type}}"] is "{{LanguageModelMessageType/tool-call}}", then set |hasToolCallInExpectedOutputs| to true. + 1. If |expected|["{{LanguageModelExpected/languages}}"] [=map/exists=], then [=validate and canonicalize language tags=] given |expected| and "{{LanguageModelExpected/languages}}". + + 1. If |options|["{{LanguageModelCreateCoreOptions/tools}}"] [=map/exists=] and |options|["{{LanguageModelCreateCoreOptions/tools}}"] is not [=list/is empty|empty=], then: + 1. If |hasToolCallInExpectedOutputs| is false, then throw a {{TypeError}}. + 1. Let |toolNames| be an empty [=ordered set=] of [=strings=]. + 1. [=list/For each=] |tool| of |options|["{{LanguageModelCreateCoreOptions/tools}}"]: + 1. If |tool|["{{LanguageModelToolDeclaration/name}}"] is the empty [=string=], then throw a {{TypeError}}. + 1. If |toolNames| [=set/contains=] |tool|["{{LanguageModelToolDeclaration/name}}"], then throw a {{TypeError}}. + 1. [=set/Append=] |tool|["{{LanguageModelToolDeclaration/name}}"] to |toolNames|. + 1. If |tool|["{{LanguageModelToolDeclaration/description}}"] is the empty [=string=], then throw a {{TypeError}}. + 1. Let |schema| be |tool|["{{LanguageModelToolDeclaration/inputSchema}}"]. + 1. Let |typeValue| be ? [$Get$](<|schema|, "`type`">). + 1. If |typeValue| is not "`object`", then throw a {{TypeError}}. + 1. Let |propertiesValue| be ? [$Get$](<|schema|, "`properties`">). + 1. If |propertiesValue| is not undefined and |propertiesValue| is not an [=Object=], then throw a {{TypeError}}. + 1. Let |requiredValue| be ? [$Get$](<|schema|, "`required`">). + 1. If |requiredValue| is not undefined and ? [$IsArray$](|requiredValue|) is false, then throw a {{TypeError}}. + 1. Perform ? [=serialize a JavaScript value to a JSON string=] given |schema|. 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=], then: 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. @@ -277,6 +297,7 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. Perform [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. </div> + <div algorithm> To <dfn>download the language model</dfn>, given a {{LanguageModelCreateCoreOptions}} |options|: @@ -334,13 +355,17 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. Let |initialMessages| be an empty [=list=] of {{LanguageModelMessage}}s. - 1. Let |initialMessagesUsage| be 0. + 1. Let |tools| be |options|["{{LanguageModelCreateCoreOptions/tools}}"] if it [=map/exists=]; otherwise an empty [=list=]. - 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=], then: + 1. Let |initialContextUsage| be 0. + + 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=] and is not [=list/empty=], then: 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |expectedInputs|. 1. Set |initialMessages| to the result of [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. - 1. Set |initialMessagesUsage| to the result of [=measure language model context usage=] given |initialMessages|, and |options|["{{LanguageModelCreateOptions/signal}}"]. + + 1. If |initialMessages| is not [=list/empty=] or |tools| is not [=list/empty=], then: + 1. Set |initialContextUsage| to the amount of context window used to encode |initialMessages| and |tools|. 1. Return a new {{LanguageModel}} object, created in |realm|, with @@ -364,13 +389,13 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe :: |options|["{{LanguageModelCreateCoreOptions/expectedOutputs}}"] if it [=map/exists=]; otherwise an empty [=list=] : [=LanguageModel/tools=] - :: |options|["{{LanguageModelCreateCoreOptions/tools}}"] if it [=map/exists=]; otherwise an empty [=list=] + :: |tools| : [=LanguageModel/context window size=] :: |contextWindowSize| : [=LanguageModel/current context usage=] - :: |initialMessagesUsage| + :: |initialContextUsage| </dl> </div> @@ -590,6 +615,8 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. [=Assert=]: this algorithm is running [=in parallel=]. + 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |model|'s [=LanguageModel/expected inputs=]. + 1. Let |messages| be the result of [=validating and canonicalizing a prompt=] given |input|, |expectedInputTypes|, and true if |model|'s [=LanguageModel/current context usage=] is greater than 0, otherwise false. If this throws an exception |e|, then: @@ -614,8 +641,6 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. Perform |error| given |errorInfo|. 1. Return false. - 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |model|'s [=LanguageModel/expected inputs=]. - 1. In an [=implementation-defined=] manner, update the underlying model's internal state to include |messages|. The process should use |model|'s [=LanguageModel/initial messages=], |model|'s [=LanguageModel/sampling mode=], |model|'s [=LanguageModel/top K=], |model|'s [=LanguageModel/temperature=], |model|'s [=LanguageModel/expected inputs=], |model|'s [=LanguageModel/expected outputs=], and |model|'s [=LanguageModel/tools=] to guide how the state is updated. @@ -639,7 +664,7 @@ The following are the [=event handlers=] (and their corresponding [=event handle * a {{LanguageModel}} |model|, * an object-or-null |responseConstraint|, - * an algorithm-or-null |chunkProduced| that takes a [=string=] and returns nothing, + * an algorithm-or-null |chunkProduced| that takes a [=string=] or a {{LanguageModelMessageContent}} and returns nothing, * an algorithm-or-null |done| that takes no arguments and returns nothing, * an algorithm-or-null |error| that takes [=error information=] and returns nothing, and * an algorithm-or-null |stopProducing| that takes no arguments and returns a boolean, @@ -654,20 +679,39 @@ The following are the [=event handlers=] (and their corresponding [=event handle The prompting process must conform to the guidance given in [[#privacy]] and [[#security]]. - If |model|'s [=LanguageModel/tools=] is not empty, the model may use the provided tools by calling their <var ignore>execute</var> functions. + If |model|'s [=LanguageModel/tools=] is not [=list/empty=], the model may produce one or more tool calls based on the declared tools in |model|'s [=LanguageModel/tools=], in addition to or instead of text. 1. While true: - 1. Wait for the next chunk of response data to be produced, for the process to finish, or for the result of calling |stopProducing| to become true. + 1. Wait for the next chunk of response data (text or a tool call) to be produced, for the process to finish, or for the result of calling |stopProducing| to become true. - 1. If such a chunk is successfully produced: + 1. If a text chunk is successfully produced: 1. Let it be represented as a [=string=] |chunk|. 1. If |chunkProduced| is not null, perform |chunkProduced| given |chunk|. + 1. Otherwise, if a tool call is successfully produced: + + 1. Let |callID| be an [=implementation-defined=] non-empty [=string=] identifying the tool call. + + 1. Let |name| be a [=string=] representing the name of the tool being called. + + 1. Let |arguments| be an [=ECMAScript/Object=] representing the JSON object of arguments for the tool call, created in |model|'s [=relevant realm=]. (If the arguments cannot be converted to an [=ECMAScript/Object=], treat this as a "{{DataError}}" failure below.) + + 1. Let |toolCall| be a new {{LanguageModelToolCall}} created in |model|'s [=relevant realm=] with [=LanguageModelToolCall/call ID=] set to |callID|, [=LanguageModelToolCall/name=] set to |name|, and [=LanguageModelToolCall/arguments=] set to |arguments|. + + 1. Let |toolCallContent| be a {{LanguageModelMessageContent}} initialized with <span style="white-space: pre-wrap">«[ + "{{LanguageModelMessageContent/type}}" → "{{LanguageModelMessageType/tool-call}}", + "{{LanguageModelMessageContent/value}}" → |toolCall| + ]»</span>. + + 1. If |chunkProduced| is not null, perform |chunkProduced| given |toolCallContent|. + 1. Otherwise, if the process has finished: + 1. In an [=implementation-defined=] manner, update the underlying model's internal state and |model|'s [=LanguageModel/current context usage=] to include the generated response (both text and any tool calls). + 1. If |done| is not null, perform |done|. 1. [=iteration/Break=]. @@ -683,6 +727,7 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. [=iteration/Break=]. </div> + <h4 id="language-model-usage">Usage</h4> <div algorithm> @@ -781,11 +826,11 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. If |message|["{{LanguageModelMessage/role}}"] is not "{{LanguageModelMessageRole/assistant}}", then throw a "{{SyntaxError}}" {{DOMException}}. - 1. If |message| is not the last item in |messages|, then throw a "{{SyntaxError}}" {{DOMException}}. + 1. If |message| is not the last item in |input|, then throw a "{{SyntaxError}}" {{DOMException}}. 1. If |message|["{{LanguageModelMessage/role}}"] is "{{LanguageModelMessageRole/system}}", then: - 1. If |hasAppendedInput| is true, then throw a "{{TypeError}}" {{DOMException}}. + 1. If |hasAppendedInput| is true, then throw a {{TypeError}}. 1. If |message|["{{LanguageModelMessage/content}}"] is an empty [=list=], then: @@ -793,26 +838,52 @@ The following are the [=event handlers=] (and their corresponding [=event handle "{{LanguageModelMessageContent/type}}" → "{{LanguageModelMessageType/text}}", "{{LanguageModelMessageContent/value}}" → "" ]»</span>. - - 1. [=list/append=] |emptyContent| to |message|["{{LanguageModelMessage/content}}"]. - + + 1. [=list/Append=] |emptyContent| to |message|["{{LanguageModelMessage/content}}"]. + 1. [=list/For each=] |content| of |message|["{{LanguageModelMessage/content}}"]: - 1. If |message|["{{LanguageModelMessage/role}}"] is "{{LanguageModelMessageRole/assistant}}" and |content|["{{LanguageModelMessageContent/type}}"] is not "{{LanguageModelMessageType/text}}", then throw a "{{NotSupportedError}}" {{DOMException}}. + 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/tool-call}}" and |message|["{{LanguageModelMessage/role}}"] is not "{{LanguageModelMessageRole/assistant}}", then throw a {{TypeError}}. + + 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/tool-response}}" and |message|["{{LanguageModelMessage/role}}"] is not "{{LanguageModelMessageRole/user}}", then throw a {{TypeError}}. + + 1. If |message|["{{LanguageModelMessage/role}}"] is "{{LanguageModelMessageRole/assistant}}" and |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/image}}" or "{{LanguageModelMessageType/audio}}", then throw a "{{NotSupportedError}}" {{DOMException}}. - 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/text}}" and |content|["{{LanguageModelMessageContent/value}}"] is not a [=string=], then throw a "{{TypeError}}" {{DOMException}}. + 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/text}}" and |content|["{{LanguageModelMessageContent/value}}"] is not a [=string=], then throw a {{TypeError}}. 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/image}}", then: 1. If |expectedTypes| does not [=list/contain=] "{{LanguageModelMessageType/image}}", then throw a "{{NotSupportedError}}" {{DOMException}}. - 1. If |content|["{{LanguageModelMessageContent/value}}"] is not an {{ImageBitmapSource}} or {{BufferSource}}, then throw a "{{TypeError}}" {{DOMException}}. + 1. If |content|["{{LanguageModelMessageContent/value}}"] is not an {{ImageBitmapSource}} or {{BufferSource}}, then throw a {{TypeError}}. 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/audio}}", then: 1. If |expectedTypes| does not [=list/contain=] "{{LanguageModelMessageType/audio}}", then throw a "{{NotSupportedError}}" {{DOMException}}. - 1. If |content|["{{LanguageModelMessageContent/value}}"] is not an {{AudioBuffer}}, {{BufferSource}}, or {{Blob}}, then throw a "{{TypeError}}" {{DOMException}}. + 1. If |content|["{{LanguageModelMessageContent/value}}"] is not an {{AudioBuffer}}, {{BufferSource}}, or {{Blob}}, then throw a {{TypeError}}. + + 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/tool-call}}", then: + + 1. If |expectedTypes| does not [=list/contain=] "{{LanguageModelMessageType/tool-call}}", then throw a "{{NotSupportedError}}" {{DOMException}}. + + 1. If |content|["{{LanguageModelMessageContent/value}}"] is not a {{LanguageModelToolCall}}, then throw a {{TypeError}}. + + 1. Let |arguments| be |content|["{{LanguageModelMessageContent/value}}"]'s [=LanguageModelToolCall/arguments=]. + + 1. If |arguments| is not null, then: + 1. If ? [$IsArray$](|arguments|) is true, or |arguments| is a [=platform object=], or |arguments| cannot be serialized to a JSON object (for example, due to circular references or non-JSON-serializable values such as functions or BigInts), then throw a "{{DataError}}" {{DOMException}}. + + 1. If |content|["{{LanguageModelMessageContent/type}}"] is "{{LanguageModelMessageType/tool-response}}", then: + + 1. If |expectedTypes| does not [=list/contain=] "{{LanguageModelMessageType/tool-response}}", then throw a "{{NotSupportedError}}" {{DOMException}}. + + 1. If |content|["{{LanguageModelMessageContent/value}}"] is not a {{LanguageModelToolResponse}}, then throw a {{TypeError}}. + + 1. If |content|["{{LanguageModelMessageContent/value}}"] is a {{LanguageModelToolSuccess}}, then [=list/for each=] |resultItem| of |content|["{{LanguageModelMessageContent/value}}"]'s [=LanguageModelToolSuccess/result=]: + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/image}}" or "{{LanguageModelToolResultType/audio}}" and the user agent does not support multimodal tool result content, then throw a "{{NotSupportedError}}" {{DOMException}}. + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/text}}" or "{{LanguageModelToolResultType/object}}", then: + 1. If |resultItem|["{{LanguageModelToolResultContent/value}}"] is a [=platform object=], or cannot be serialized to JSON (for example, due to circular references or non-JSON-serializable values such as functions or BigInts), then throw a "{{DataError}}" {{DOMException}}. 1. Let |contentWithContiguousTextCollapsed| be an empty [=list=] of {{LanguageModelMessageContent}}s. @@ -942,6 +1013,79 @@ When prompting fails, the following possible reasons may be surfaced to the web 1. Return |promise|. </div> +<h3 id="the-languagemodeltoolcall-class">The {{LanguageModelToolCall}} class</h3> + +Every {{LanguageModelToolCall}} has a <dfn for="LanguageModelToolCall">call ID</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolCall}} has a <dfn for="LanguageModelToolCall">name</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolCall}} has an <dfn for="LanguageModelToolCall">arguments</dfn>, an [=ECMAScript/object=] or null, set during creation. + +<div algorithm> + The <dfn constructor for="LanguageModelToolCall" lt="LanguageModelToolCall(init)">new LanguageModelToolCall(|init|)</dfn> constructor steps are: + + 1. Set [=this=]'s [=LanguageModelToolCall/call ID=] to |init|["{{LanguageModelToolCallInit/callID}}"]. + + 1. Set [=this=]'s [=LanguageModelToolCall/name=] to |init|["{{LanguageModelToolCallInit/name}}"]. + + 1. Set [=this=]'s [=LanguageModelToolCall/arguments=] to |init|["{{LanguageModelToolCallInit/arguments}}"] if it [=map/exists=]; otherwise null. +</div> + +The <dfn attribute for="LanguageModelToolCall">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/call ID=]. + +The <dfn attribute for="LanguageModelToolCall">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/name=]. + +The <dfn attribute for="LanguageModelToolCall">arguments</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/arguments=]. + +<h3 id="the-languagemodeltoolsuccess-class">The {{LanguageModelToolSuccess}} class</h3> + +Every {{LanguageModelToolSuccess}} has a <dfn for="LanguageModelToolSuccess">call ID</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolSuccess}} has a <dfn for="LanguageModelToolSuccess">name</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolSuccess}} has a <dfn for="LanguageModelToolSuccess">result</dfn>, a <code>{{FrozenArray}}&lt;{{LanguageModelToolResultContent}}></code>, set during creation. + +<div algorithm> + The <dfn constructor for="LanguageModelToolSuccess" lt="LanguageModelToolSuccess(init)">new LanguageModelToolSuccess(|init|)</dfn> constructor steps are: + + 1. Set [=this=]'s [=LanguageModelToolSuccess/call ID=] to |init|["{{LanguageModelToolSuccessInit/callID}}"]. + + 1. Set [=this=]'s [=LanguageModelToolSuccess/name=] to |init|["{{LanguageModelToolSuccessInit/name}}"]. + + 1. Set [=this=]'s [=LanguageModelToolSuccess/result=] to the result of [=creating a frozen array=] from |init|["{{LanguageModelToolSuccessInit/result}}"]. +</div> + +The <dfn attribute for="LanguageModelToolSuccess">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/call ID=]. + +The <dfn attribute for="LanguageModelToolSuccess">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/name=]. + +The <dfn attribute for="LanguageModelToolSuccess">result</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/result=]. + +<h3 id="the-languagemodeltoolerror-class">The {{LanguageModelToolError}} class</h3> + +Every {{LanguageModelToolError}} has a <dfn for="LanguageModelToolError">call ID</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolError}} has a <dfn for="LanguageModelToolError">name</dfn>, a [=string=], set during creation. + +Every {{LanguageModelToolError}} has an <dfn for="LanguageModelToolError">error message</dfn>, a [=string=], set during creation. + +<div algorithm> + The <dfn constructor for="LanguageModelToolError" lt="LanguageModelToolError(init)">new LanguageModelToolError(|init|)</dfn> constructor steps are: + + 1. Set [=this=]'s [=LanguageModelToolError/call ID=] to |init|["{{LanguageModelToolErrorInit/callID}}"]. + + 1. Set [=this=]'s [=LanguageModelToolError/name=] to |init|["{{LanguageModelToolErrorInit/name}}"]. + + 1. Set [=this=]'s [=LanguageModelToolError/error message=] to |init|["{{LanguageModelToolErrorInit/errorMessage}}"]. +</div> + +The <dfn attribute for="LanguageModelToolError">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/call ID=]. + +The <dfn attribute for="LanguageModelToolError">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/name=]. + +The <dfn attribute for="LanguageModelToolError">errorMessage</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/error message=]. + + <h3 id="permissions-policy">Permissions policy integration</h3> Access to the prompt API is gated behind the [=policy-controlled feature=] "<dfn permission>language-model</dfn>", which has a [=policy-controlled feature/default allowlist=] of <code>[=default allowlist/'self'=]</code>. From 3d7b68b22d0c713b395c8afdeca31f893a4d1f1a Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Tue, 29 Sep 2026 14:17:11 -0700 Subject: [PATCH 17/30] add note to spec --- index.bs | 3 +++ 1 file changed, 3 insertions(+) diff --git a/index.bs b/index.bs index e289ba6..a0769ff 100644 --- a/index.bs +++ b/index.bs @@ -108,6 +108,9 @@ interface LanguageModelParams { }; // A description of a tool call that a language model can invoke. +// Note: When considering changes to this dictionary, authors should ensure +// general alignment with ModelContextTool from WebMCP +// (https://webmachinelearning.github.io/webmcp/#model-context-tool). dictionary LanguageModelToolDeclaration { required DOMString name; required DOMString description; From 4a192b1048f9ec9e5a71d386dfcfdd724f357015 Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Wed, 30 Sep 2026 16:10:07 -0700 Subject: [PATCH 18/30] rename callID and add defaults in spec --- index.bs | 44 ++++++++++++++++++++++---------------------- 1 file changed, 22 insertions(+), 22 deletions(-) diff --git a/index.bs b/index.bs index a0769ff..e4fdb49 100644 --- a/index.bs +++ b/index.bs @@ -135,14 +135,14 @@ dictionary LanguageModelCreateCoreOptions { // Tools that the language model can use. // **EXPERIMENTAL**: Only available in experimental contexts. - sequence<LanguageModelToolDeclaration> tools; + sequence<LanguageModelToolDeclaration> tools = []; }; dictionary LanguageModelCreateOptions : LanguageModelCreateCoreOptions { AbortSignal signal; CreateMonitorCallback monitor; - sequence<LanguageModelMessage> initialPrompts; + sequence<LanguageModelMessage> initialPrompts = []; }; dictionary LanguageModelPromptOptions { @@ -214,37 +214,37 @@ dictionary LanguageModelToolResultContent { [Exposed=Window, SecureContext] interface LanguageModelToolCall { constructor(LanguageModelToolCallInit init); - readonly attribute DOMString callID; + readonly attribute DOMString callId; readonly attribute DOMString name; readonly attribute object? arguments; }; dictionary LanguageModelToolCallInit { - required DOMString callID; + required DOMString callId; required DOMString name; object arguments; }; [Exposed=Window, SecureContext] interface LanguageModelToolSuccess { constructor(LanguageModelToolSuccessInit init); - readonly attribute DOMString callID; + readonly attribute DOMString callId; readonly attribute DOMString name; readonly attribute FrozenArray<LanguageModelToolResultContent> result; }; dictionary LanguageModelToolSuccessInit { - required DOMString callID; + required DOMString callId; required DOMString name; required sequence<LanguageModelToolResultContent> result; }; [Exposed=Window, SecureContext] interface LanguageModelToolError { constructor(LanguageModelToolErrorInit init); - readonly attribute DOMString callID; + readonly attribute DOMString callId; readonly attribute DOMString name; readonly attribute DOMString errorMessage; }; dictionary LanguageModelToolErrorInit { - required DOMString callID; + required DOMString callId; required DOMString name; required DOMString errorMessage; }; @@ -277,7 +277,7 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. If |expected|["{{LanguageModelExpected/type}}"] is "{{LanguageModelMessageType/tool-call}}", then set |hasToolCallInExpectedOutputs| to true. 1. If |expected|["{{LanguageModelExpected/languages}}"] [=map/exists=], then [=validate and canonicalize language tags=] given |expected| and "{{LanguageModelExpected/languages}}". - 1. If |options|["{{LanguageModelCreateCoreOptions/tools}}"] [=map/exists=] and |options|["{{LanguageModelCreateCoreOptions/tools}}"] is not [=list/is empty|empty=], then: + 1. If |options|["{{LanguageModelCreateCoreOptions/tools}}"] is not [=list/is empty|empty=], then: 1. If |hasToolCallInExpectedOutputs| is false, then throw a {{TypeError}}. 1. Let |toolNames| be an empty [=ordered set=] of [=strings=]. 1. [=list/For each=] |tool| of |options|["{{LanguageModelCreateCoreOptions/tools}}"]: @@ -294,7 +294,7 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. If |requiredValue| is not undefined and ? [$IsArray$](|requiredValue|) is false, then throw a {{TypeError}}. 1. Perform ? [=serialize a JavaScript value to a JSON string=] given |schema|. - 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=], then: + 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=] and is not [=list/is empty|empty=], then: 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |expectedInputs|. 1. Perform [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. @@ -326,13 +326,13 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe This could include loading the appropriate model and any fine-tunings necessary to support |options| into memory. - 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=], then: + 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] is not [=list/is empty|empty=], then: 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |expectedInputs|. 1. Let |initialMessages| be the result of [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. 1. Load |initialMessages| into the model's context window. - 1. If |options|["{{LanguageModelCreateCoreOptions/tools}}"] [=map/exists=], then load |options|["{{LanguageModelCreateCoreOptions/tools}}"] into the model's context window. + 1. If |options|["{{LanguageModelCreateCoreOptions/tools}}"] is not [=list/is empty|empty=], then load |options|["{{LanguageModelCreateCoreOptions/tools}}"] into the model's context window. 1. If initialization failed because the process of loading |options| resulted in using up all of the model's context window, then: @@ -358,12 +358,12 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. Let |initialMessages| be an empty [=list=] of {{LanguageModelMessage}}s. - 1. Let |tools| be |options|["{{LanguageModelCreateCoreOptions/tools}}"] if it [=map/exists=]; otherwise an empty [=list=]. + 1. Let |tools| be |options|["{{LanguageModelCreateCoreOptions/tools}}"]. 1. Let |initialContextUsage| be 0. 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=] and is not [=list/empty=], then: - 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. + 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"]. 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |expectedInputs|. 1. Set |initialMessages| to the result of [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. @@ -696,13 +696,13 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. Otherwise, if a tool call is successfully produced: - 1. Let |callID| be an [=implementation-defined=] non-empty [=string=] identifying the tool call. + 1. Let |callId| be an [=implementation-defined=] non-empty [=string=] identifying the tool call. 1. Let |name| be a [=string=] representing the name of the tool being called. 1. Let |arguments| be an [=ECMAScript/Object=] representing the JSON object of arguments for the tool call, created in |model|'s [=relevant realm=]. (If the arguments cannot be converted to an [=ECMAScript/Object=], treat this as a "{{DataError}}" failure below.) - 1. Let |toolCall| be a new {{LanguageModelToolCall}} created in |model|'s [=relevant realm=] with [=LanguageModelToolCall/call ID=] set to |callID|, [=LanguageModelToolCall/name=] set to |name|, and [=LanguageModelToolCall/arguments=] set to |arguments|. + 1. Let |toolCall| be a new {{LanguageModelToolCall}} created in |model|'s [=relevant realm=] with [=LanguageModelToolCall/call ID=] set to |callId|, [=LanguageModelToolCall/name=] set to |name|, and [=LanguageModelToolCall/arguments=] set to |arguments|. 1. Let |toolCallContent| be a {{LanguageModelMessageContent}} initialized with <span style="white-space: pre-wrap">«[ "{{LanguageModelMessageContent/type}}" → "{{LanguageModelMessageType/tool-call}}", @@ -1027,14 +1027,14 @@ Every {{LanguageModelToolCall}} has an <dfn for="LanguageModelToolCall">argument <div algorithm> The <dfn constructor for="LanguageModelToolCall" lt="LanguageModelToolCall(init)">new LanguageModelToolCall(|init|)</dfn> constructor steps are: - 1. Set [=this=]'s [=LanguageModelToolCall/call ID=] to |init|["{{LanguageModelToolCallInit/callID}}"]. + 1. Set [=this=]'s [=LanguageModelToolCall/call ID=] to |init|["{{LanguageModelToolCallInit/callId}}"]. 1. Set [=this=]'s [=LanguageModelToolCall/name=] to |init|["{{LanguageModelToolCallInit/name}}"]. 1. Set [=this=]'s [=LanguageModelToolCall/arguments=] to |init|["{{LanguageModelToolCallInit/arguments}}"] if it [=map/exists=]; otherwise null. </div> -The <dfn attribute for="LanguageModelToolCall">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/call ID=]. +The <dfn attribute for="LanguageModelToolCall">callId</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/call ID=]. The <dfn attribute for="LanguageModelToolCall">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolCall/name=]. @@ -1051,14 +1051,14 @@ Every {{LanguageModelToolSuccess}} has a <dfn for="LanguageModelToolSuccess">res <div algorithm> The <dfn constructor for="LanguageModelToolSuccess" lt="LanguageModelToolSuccess(init)">new LanguageModelToolSuccess(|init|)</dfn> constructor steps are: - 1. Set [=this=]'s [=LanguageModelToolSuccess/call ID=] to |init|["{{LanguageModelToolSuccessInit/callID}}"]. + 1. Set [=this=]'s [=LanguageModelToolSuccess/call ID=] to |init|["{{LanguageModelToolSuccessInit/callId}}"]. 1. Set [=this=]'s [=LanguageModelToolSuccess/name=] to |init|["{{LanguageModelToolSuccessInit/name}}"]. 1. Set [=this=]'s [=LanguageModelToolSuccess/result=] to the result of [=creating a frozen array=] from |init|["{{LanguageModelToolSuccessInit/result}}"]. </div> -The <dfn attribute for="LanguageModelToolSuccess">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/call ID=]. +The <dfn attribute for="LanguageModelToolSuccess">callId</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/call ID=]. The <dfn attribute for="LanguageModelToolSuccess">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolSuccess/name=]. @@ -1075,14 +1075,14 @@ Every {{LanguageModelToolError}} has an <dfn for="LanguageModelToolError">error <div algorithm> The <dfn constructor for="LanguageModelToolError" lt="LanguageModelToolError(init)">new LanguageModelToolError(|init|)</dfn> constructor steps are: - 1. Set [=this=]'s [=LanguageModelToolError/call ID=] to |init|["{{LanguageModelToolErrorInit/callID}}"]. + 1. Set [=this=]'s [=LanguageModelToolError/call ID=] to |init|["{{LanguageModelToolErrorInit/callId}}"]. 1. Set [=this=]'s [=LanguageModelToolError/name=] to |init|["{{LanguageModelToolErrorInit/name}}"]. 1. Set [=this=]'s [=LanguageModelToolError/error message=] to |init|["{{LanguageModelToolErrorInit/errorMessage}}"]. </div> -The <dfn attribute for="LanguageModelToolError">callID</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/call ID=]. +The <dfn attribute for="LanguageModelToolError">callId</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/call ID=]. The <dfn attribute for="LanguageModelToolError">name</dfn> getter steps are to return [=this=]'s [=LanguageModelToolError/name=]. From 9730635fcaca581df14cd5a7a19e32a79ccdc10f Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Wed, 30 Sep 2026 16:55:56 -0700 Subject: [PATCH 19/30] update explainer to address comments --- README.md | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index e604041..33f17f4 100644 --- a/README.md +++ b/README.md @@ -221,7 +221,9 @@ await session.append([ { type: "tool-call", value: new LanguageModelToolCall({ - callID: "get_weather_1", + // In few-shot examples, `callID` can be any string as + // long as the corresponding tool-response uses the same `callID`. + callID: "example-call-1", name: "get_weather", arguments: { location: "Seattle" }, }), @@ -234,7 +236,7 @@ await session.append([ { type: "tool-response", value: new LanguageModelToolSuccess({ - callID: "get_weather_1", + callID: "example-call-1", name: "get_weather", result: [ { type: "object", value: { temperature: "55F", humidity: "67%" } }, @@ -263,7 +265,7 @@ Open loop is enabled by specifying `{ type: "tool-call" }` in `expectedOutputs` When the model does not invoke any tools, `session.prompt()` resolves to a `DOMString` as usual. When a tool needs to be called, `session.prompt()` resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence<LanguageModelMessageContent>`). If the model outputs both text and tool calls, it's resolved to an array, where the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. Each `LanguageModelToolCall` object contains: -* `callID`: A unique string identifier for this tool call. +* `callID`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callID` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. * `name`: The name of the tool to invoke. * `arguments`: A dictionary fitting the JSON `inputSchema` of the tool's declaration (which must have `type: "object"`). @@ -281,6 +283,9 @@ if (Array.isArray(result)) { if (toolCallMsg && toolCallMsg.value.name === "get_weather") { const toolCall = toolCallMsg.value; const toolResult = getWeather(toolCall.arguments.location); + // For simplicity, this example assumes a single tool call followed by a + // final text response. In practice, the model may respond with additional + // tool calls, which would typically be handled in a loop. result = await session.prompt([ { role: "user", From 1477f4832b40f4f1e2edf8cf1436aac74580f7ae Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Wed, 30 Sep 2026 16:57:00 -0700 Subject: [PATCH 20/30] Rename callID to callId in explainer to match spec --- README.md | 16 ++++++++-------- 1 file changed, 8 insertions(+), 8 deletions(-) diff --git a/README.md b/README.md index 33f17f4..5b83260 100644 --- a/README.md +++ b/README.md @@ -221,9 +221,9 @@ await session.append([ { type: "tool-call", value: new LanguageModelToolCall({ - // In few-shot examples, `callID` can be any string as - // long as the corresponding tool-response uses the same `callID`. - callID: "example-call-1", + // In few-shot examples, `callId` can be any string as + // long as the corresponding tool-response uses the same `callId`. + callId: "example-call-1", name: "get_weather", arguments: { location: "Seattle" }, }), @@ -236,7 +236,7 @@ await session.append([ { type: "tool-response", value: new LanguageModelToolSuccess({ - callID: "example-call-1", + callId: "example-call-1", name: "get_weather", result: [ { type: "object", value: { temperature: "55F", humidity: "67%" } }, @@ -254,8 +254,8 @@ await session.append([ Note that: * Message `content` `type` supports `"tool-call"` and `"tool-response"`: - * `"tool-call"` content must use `role: "assistant"` and its `value` must be a `LanguageModelToolCall` instance (`new LanguageModelToolCall({ callID, name, arguments })`). - * `"tool-response"` content must use `role: "user"` and its `value` must be either a `LanguageModelToolSuccess` instance (`new LanguageModelToolSuccess({ callID, name, result })`) or a `LanguageModelToolError` instance (`new LanguageModelToolError({ callID, name, errorMessage })`). + * `"tool-call"` content must use `role: "assistant"` and its `value` must be a `LanguageModelToolCall` instance (`new LanguageModelToolCall({ callId, name, arguments })`). + * `"tool-response"` content must use `role: "user"` and its `value` must be either a `LanguageModelToolSuccess` instance (`new LanguageModelToolSuccess({ callId, name, result })`) or a `LanguageModelToolError` instance (`new LanguageModelToolError({ callId, name, errorMessage })`). * `LanguageModelToolSuccess.result` is a list of `{ type, value }` dictionaries, where `type` can be `"text"`, `"image"`, `"audio"`, or `"object"`, and `value` is `any`. #### Open Loop @@ -265,7 +265,7 @@ Open loop is enabled by specifying `{ type: "tool-call" }` in `expectedOutputs` When the model does not invoke any tools, `session.prompt()` resolves to a `DOMString` as usual. When a tool needs to be called, `session.prompt()` resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence<LanguageModelMessageContent>`). If the model outputs both text and tool calls, it's resolved to an array, where the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. Each `LanguageModelToolCall` object contains: -* `callID`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callID` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. +* `callId`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callId` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. * `name`: The name of the tool to invoke. * `arguments`: A dictionary fitting the JSON `inputSchema` of the tool's declaration (which must have `type: "object"`). @@ -293,7 +293,7 @@ if (Array.isArray(result)) { { type: "tool-response", value: new LanguageModelToolSuccess({ - callID: toolCall.callID, + callId: toolCall.callId, name: toolCall.name, result: [{ type: "object", value: toolResult }], }), From 268e3d1103d5ca98751e455b82fb831bf342e6f7 Mon Sep 17 00:00:00 2001 From: jingyun19 <jingyun@google.com> Date: Thu, 1 Oct 2026 09:39:06 -0700 Subject: [PATCH 21/30] update spec to address misc formatting issues --- index.bs | 21 +++++++++++---------- 1 file changed, 11 insertions(+), 10 deletions(-) diff --git a/index.bs b/index.bs index e4fdb49..5147fd9 100644 --- a/index.bs +++ b/index.bs @@ -52,7 +52,7 @@ interface LanguageModel : EventTarget { // **EXPERIMENTAL**: Only available in extension and experimental contexts. static Promise<LanguageModelParams?> params(); - // These will throw "TypeError" DOMExceptions if role = "system" + // These will throw a TypeError if role = "system" Promise<LanguageModelPromptResult> prompt( LanguageModelPrompt input, optional LanguageModelPromptOptions options = {} @@ -186,7 +186,6 @@ dictionary LanguageModelMessageContent { required LanguageModelMessageValue value; }; - enum LanguageModelSamplingMode { "most-predictable", "predictable", "slightly-predictable", "balanced", "slightly-creative", "creative", "most-creative" }; enum LanguageModelMessageRole { "system", "user", "assistant" }; @@ -224,6 +223,7 @@ dictionary LanguageModelToolCallInit { required DOMString name; object arguments; }; + [Exposed=Window, SecureContext] interface LanguageModelToolSuccess { constructor(LanguageModelToolSuccessInit init); @@ -231,11 +231,13 @@ interface LanguageModelToolSuccess { readonly attribute DOMString name; readonly attribute FrozenArray<LanguageModelToolResultContent> result; }; + dictionary LanguageModelToolSuccessInit { required DOMString callId; required DOMString name; required sequence<LanguageModelToolResultContent> result; }; + [Exposed=Window, SecureContext] interface LanguageModelToolError { constructor(LanguageModelToolErrorInit init); @@ -243,6 +245,7 @@ interface LanguageModelToolError { readonly attribute DOMString name; readonly attribute DOMString errorMessage; }; + dictionary LanguageModelToolErrorInit { required DOMString callId; required DOMString name; @@ -252,7 +255,6 @@ dictionary LanguageModelToolErrorInit { // The response from executing a tool call - either success or error. typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolResponse; -

Creation

@@ -300,7 +302,6 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. Perform [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. -
To download the language model, given a {{LanguageModelCreateCoreOptions}} |options|: @@ -362,12 +363,12 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. Let |initialContextUsage| be 0. - 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=] and is not [=list/empty=], then: - 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"]. + 1. If |options|["{{LanguageModelCreateOptions/initialPrompts}}"] [=map/exists=] and is not [=list/is empty|empty=], then: + 1. Let |expectedInputs| be |options|["{{LanguageModelCreateCoreOptions/expectedInputs}}"] if it [=map/exists=]; otherwise an empty [=list=]. 1. Let |expectedInputTypes| be the result of [=get the expected content types=] given |expectedInputs|. 1. Set |initialMessages| to the result of [=validating and canonicalizing a prompt=] given |options|["{{LanguageModelCreateOptions/initialPrompts}}"], |expectedInputTypes|, and false. - 1. If |initialMessages| is not [=list/empty=] or |tools| is not [=list/empty=], then: + 1. If |initialMessages| is not [=list/is empty|empty=] or |tools| is not [=list/is empty|empty=], then: 1. Set |initialContextUsage| to the amount of context window used to encode |initialMessages| and |tools|. 1. Return a new {{LanguageModel}} object, created in |realm|, with @@ -682,7 +683,7 @@ The following are the [=event handlers=] (and their corresponding [=event handle The prompting process must conform to the guidance given in [[#privacy]] and [[#security]]. - If |model|'s [=LanguageModel/tools=] is not [=list/empty=], the model may produce one or more tool calls based on the declared tools in |model|'s [=LanguageModel/tools=], in addition to or instead of text. + If |model|'s [=LanguageModel/tools=] is not [=list/is empty|empty=], the model may produce one or more tool calls based on the declared tools in |model|'s [=LanguageModel/tools=], in addition to or instead of text. 1. While true: @@ -700,7 +701,7 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. Let |name| be a [=string=] representing the name of the tool being called. - 1. Let |arguments| be an [=ECMAScript/Object=] representing the JSON object of arguments for the tool call, created in |model|'s [=relevant realm=]. (If the arguments cannot be converted to an [=ECMAScript/Object=], treat this as a "{{DataError}}" failure below.) + 1. Let |arguments| be an [=Object=] representing the JSON object of arguments for the tool call, created in |model|'s [=relevant realm=]. 1. Let |toolCall| be a new {{LanguageModelToolCall}} created in |model|'s [=relevant realm=] with [=LanguageModelToolCall/call ID=] set to |callId|, [=LanguageModelToolCall/name=] set to |name|, and [=LanguageModelToolCall/arguments=] set to |arguments|. @@ -1046,7 +1047,7 @@ Every {{LanguageModelToolSuccess}} has a cal Every {{LanguageModelToolSuccess}} has a name, a [=string=], set during creation. -Every {{LanguageModelToolSuccess}} has a result, a {{FrozenArray}}<{{LanguageModelToolResultContent}}>, set during creation. +Every {{LanguageModelToolSuccess}} has a result, a [=frozen array=] of {{LanguageModelToolResultContent}}s, set during creation.
The new LanguageModelToolSuccess(|init|) constructor steps are: From 81ff758b1b4a5c18652aab38de5dacedf226fc88 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:14:23 -0700 Subject: [PATCH 22/30] Refactor type and properties retrieval in index.bs --- index.bs | 12 ++++++++---- 1 file changed, 8 insertions(+), 4 deletions(-) diff --git a/index.bs b/index.bs index 5147fd9..cdce0a1 100644 --- a/index.bs +++ b/index.bs @@ -288,11 +288,11 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. [=set/Append=] |tool|["{{LanguageModelToolDeclaration/name}}"] to |toolNames|. 1. If |tool|["{{LanguageModelToolDeclaration/description}}"] is the empty [=string=], then throw a {{TypeError}}. 1. Let |schema| be |tool|["{{LanguageModelToolDeclaration/inputSchema}}"]. - 1. Let |typeValue| be ? [$Get$](<|schema|, "`type`">). + 1. Let |typeValue| be ? [$Get$](|schema|, "`type`"). 1. If |typeValue| is not "`object`", then throw a {{TypeError}}. - 1. Let |propertiesValue| be ? [$Get$](<|schema|, "`properties`">). + 1. Let |propertiesValue| be ? [$Get$](|schema|, "`properties`"). 1. If |propertiesValue| is not undefined and |propertiesValue| is not an [=Object=], then throw a {{TypeError}}. - 1. Let |requiredValue| be ? [$Get$](<|schema|, "`required`">). + 1. Let |requiredValue| be ? [$Get$](|schema|, "`required`"). 1. If |requiredValue| is not undefined and ? [$IsArray$](|requiredValue|) is false, then throw a {{TypeError}}. 1. Perform ? [=serialize a JavaScript value to a JSON string=] given |schema|. @@ -886,7 +886,11 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. If |content|["{{LanguageModelMessageContent/value}}"] is a {{LanguageModelToolSuccess}}, then [=list/for each=] |resultItem| of |content|["{{LanguageModelMessageContent/value}}"]'s [=LanguageModelToolSuccess/result=]: 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/image}}" or "{{LanguageModelToolResultType/audio}}" and the user agent does not support multimodal tool result content, then throw a "{{NotSupportedError}}" {{DOMException}}. - 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/text}}" or "{{LanguageModelToolResultType/object}}", then: + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/text}}" and |resultItem|["{{LanguageModelToolResultContent/value}}"] is not a [=string=], then throw a {{TypeError}}. + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/image}}" and |resultItem|["{{LanguageModelToolResultContent/value}}"] is not an {{ImageBitmapSource}} or {{BufferSource}}, then throw a {{TypeError}}. + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/audio}}" and |resultItem|["{{LanguageModelToolResultContent/value}}"] is not an {{AudioBuffer}}, {{BufferSource}}, or {{Blob}}, then throw a {{TypeError}}. + 1. If |resultItem|["{{LanguageModelToolResultContent/type}}"] is "{{LanguageModelToolResultType/object}}", then: + 1. If |resultItem|["{{LanguageModelToolResultContent/value}}"] is not an [=Object=], then throw a {{TypeError}}. 1. If |resultItem|["{{LanguageModelToolResultContent/value}}"] is a [=platform object=], or cannot be serialized to JSON (for example, due to circular references or non-JSON-serializable values such as functions or BigInts), then throw a "{{DataError}}" {{DOMException}}. 1. Let |contentWithContiguousTextCollapsed| be an empty [=list=] of {{LanguageModelMessageContent}}s. From 04f5d0a2b8d4e88fb23bd356c25169067b4abb53 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:21:21 -0700 Subject: [PATCH 23/30] Update explainer for tool use Clarified expectedInputs and LanguageModelToolSuccess.result structure in README. --- README.md | 11 ++++++++--- 1 file changed, 8 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 5b83260..7a6f7cb 100644 --- a/README.md +++ b/README.md @@ -187,7 +187,7 @@ const session = await LanguageModel.create({ content: "You are a helpful assistant. You can use tools to help the user.", }, ], - expectedInputs: [{ type: "tool-call" }, { type: "tool-response" }], + expectedInputs: [{ type: "tool-response" }], expectedOutputs: [{ type: "tool-call" }], tools: [ { @@ -208,7 +208,7 @@ const session = await LanguageModel.create({ }); ``` -In this example, the `tools` array defines a `get_weather` tool, specifying its name, description, and input schema. When `tools` are provided, `expectedOutputs` must include `{ type: "tool-call" }`. To supply tool calls and/or tool responses in `initialPrompts`, `append()`, or `prompt()`, `expectedInputs` must also include `{ type: "tool-call" }` and/or `{ type: "tool-response" }`. +In this example, the `tools` array defines a `get_weather` tool, specifying its name, description, and input schema. When `tools` are provided, `expectedOutputs` must include `{ type: "tool-call" }`, and `expectedInputs` must include `{ type: "tool-response" }` so the application can send tool execution results back to the model. (`expectedInputs` only needs to include `{ type: "tool-call" }` if the application intends to pass assistant-role tool calls as input in `initialPrompts`, `append()`, or `prompt()`, such as to provide few-shot examples or restore a saved session's context.) Few-shot examples of tool use can be appended like so: @@ -256,7 +256,12 @@ Note that: * Message `content` `type` supports `"tool-call"` and `"tool-response"`: * `"tool-call"` content must use `role: "assistant"` and its `value` must be a `LanguageModelToolCall` instance (`new LanguageModelToolCall({ callId, name, arguments })`). * `"tool-response"` content must use `role: "user"` and its `value` must be either a `LanguageModelToolSuccess` instance (`new LanguageModelToolSuccess({ callId, name, result })`) or a `LanguageModelToolError` instance (`new LanguageModelToolError({ callId, name, errorMessage })`). -* `LanguageModelToolSuccess.result` is a list of `{ type, value }` dictionaries, where `type` can be `"text"`, `"image"`, `"audio"`, or `"object"`, and `value` is `any`. +* `LanguageModelToolSuccess.result` is a list of `{ type, value }` dictionaries (`LanguageModelToolResultContent`), where `value` must match the stated `type`: + * `"text"`: a `DOMString`. + * `"image"`: an `ImageBitmapSource` or `BufferSource`. + * `"audio"`: an `AudioBuffer`, `BufferSource`, or `Blob`. + * `"object"`: a JSON-serializable JavaScript object. + #### Open Loop From 9547fe5e1e1f643a3ee22269ecd6e59e102a0c73 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:27:27 -0700 Subject: [PATCH 24/30] Enhance README with details on session.prompt() Clarify behavior of session.prompt() with expectedOutputs. --- README.md | 4 +++- 1 file changed, 3 insertions(+), 1 deletion(-) diff --git a/README.md b/README.md index 7a6f7cb..ee3156f 100644 --- a/README.md +++ b/README.md @@ -267,7 +267,9 @@ Note that: Open loop is enabled by specifying `{ type: "tool-call" }` in `expectedOutputs` when the session is created. -When the model does not invoke any tools, `session.prompt()` resolves to a `DOMString` as usual. When a tool needs to be called, `session.prompt()` resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence`). If the model outputs both text and tool calls, it's resolved to an array, where the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. +When a session is configured with only `"text"` in `expectedOutputs` (or when `expectedOutputs` is omitted), `session.prompt()` resolves to a `DOMString` as usual. + +When `expectedOutputs` includes non-text types such as `{ type: "tool-call" }`, `session.prompt()` consistently resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence`), even if the model only produces text for a given turn (e.g., `[{ type: "text", value: "..." }]`). If the model outputs both text and tool calls, the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. Each `LanguageModelToolCall` object contains: * `callId`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callId` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. From 200a2da071019dc35ca43929c662bf1ed5b1dfcf Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:32:24 -0700 Subject: [PATCH 25/30] Refactor algorithm for handling expected outputs --- index.bs | 44 ++++++++++++++++++++++++++++++++++++++++++-- 1 file changed, 42 insertions(+), 2 deletions(-) diff --git a/index.bs b/index.bs index cdce0a1..fe07b83 100644 --- a/index.bs +++ b/index.bs @@ -536,13 +536,53 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. Let |omitResponseConstraintInput| be |options|["{{LanguageModelPromptOptions/omitResponseConstraintInput}}"]. + 1. Let |hasNonTextExpectedOutput| be false. + + 1. [=list/For each=] |expected| of [=this=]'s [=LanguageModel/expected outputs=]: + 1. If |expected|["{{LanguageModelExpected/type}}"] is not "{{LanguageModelMessageType/text}}", then set |hasNonTextExpectedOutput| to true. + + 1. Let |contents| be an empty [=list=] of {{LanguageModelMessageContent}}s. + + 1. Let |text| be the empty [=string=]. + 1. Let |operation| be an algorithm step which takes arguments |chunkProduced|, |done|, |error|, and |stopProducing|, and performs the following steps: 1. Let |prefillSuccess| be the result of [=prefilling=] given [=this=], |input|, |omitResponseConstraintInput|, |responseConstraint|, |error|, and |stopProducing|. - 1. If |prefillSuccess| is true, then [=generate=] given [=this=], |responseConstraint|, |chunkProduced|, |done|, |error|, and |stopProducing|. + 1. If |prefillSuccess| is false, then return. - 1. Return the result of [=getting an aggregated AI model result=] given [=this=], |options|, and |operation|. + 1. If |hasNonTextExpectedOutput| is false: + 1. [=Generate=] given [=this=], |responseConstraint|, |chunkProduced|, |done|, |error|, and |stopProducing|. + 1. Return. + + 1. Let |onChunk| be an algorithm step which takes argument |chunk| and performs the following steps: + 1. If |chunk| is a [=string=]: + 1. Set |text| to the concatenation of |text| and |chunk|. + 1. Otherwise: + 1. [=Assert=]: |chunk| is a {{LanguageModelMessageContent}}. + 1. If |text| is not the empty [=string=]: + 1. [=list/Append=] «[ + "{{LanguageModelMessageContent/type}}" → "{{LanguageModelMessageType/text}}", + "{{LanguageModelMessageContent/value}}" → |text| + ]» to |contents|. + 1. Set |text| to the empty [=string=]. + 1. [=list/Append=] |chunk| to |contents|. + + 1. Let |onDone| be an algorithm step which takes no arguments and performs the following steps: + 1. If |text| is not the empty [=string=] or |contents| [=list/is empty=]: + 1. [=list/Append=] «[ + "{{LanguageModelMessageContent/type}}" → "{{LanguageModelMessageType/text}}", + "{{LanguageModelMessageContent/value}}" → |text| + ]» to |contents|. + 1. If |done| is not null, then perform |done|. + + 1. [=Generate=] given [=this=], |responseConstraint|, |onChunk|, |onDone|, |error|, and |stopProducing|. + + 1. Let |promise| be the result of [=getting an aggregated AI model result=] given [=this=], |options|, and |operation|. + + 1. If |hasNonTextExpectedOutput| is false, then return |promise|. + + 1. Return the result of [=reacting=] to |promise| with a fulfillment handler that returns |contents|.
From 55b4c65ebb0cc822bf70bce375ae73c2352e0b39 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:42:52 -0700 Subject: [PATCH 26/30] Clarify tool use modes and examples in README --- README.md | 170 ++++++++++++++++++++++-------------------------------- 1 file changed, 69 insertions(+), 101 deletions(-) diff --git a/README.md b/README.md index ee3156f..b8871e1 100644 --- a/README.md +++ b/README.md @@ -175,7 +175,7 @@ Because of their special behavior of being preserved on context window overflow, The Prompt API supports **tool use** via the `tools` option, allowing you to define external capabilities that a language model can invoke in a model-agnostic way. Each tool is declared with a `name`, `description`, and `inputSchema` (a JSON Schema object with `type: "object"`). -There are two tool use modes: without automatic execution (**open loop**) and with automatic execution (**closed loop**, planned). In open loop mode, the model returns tool call requests to the application, which executes the tool and sends the response back to the session. In closed loop mode, each tool declaration also includes an `execute` function that the user agent invokes automatically when the model requests a tool call. +Tool use currently operates in **open loop** mode: when the model decides to invoke a tool, it returns tool call requests (`LanguageModelToolCall`) to the application, which executes the tool and sends the result (`LanguageModelToolSuccess` or `LanguageModelToolError`) back to the session in a follow-up prompt. (See [Future work: Automatic tool execution (Closed loop)](#future-work-automatic-tool-execution-closed-loop) below for planned automatic execution support.) Here is an example of creating a session with tool declarations: @@ -210,10 +210,65 @@ const session = await LanguageModel.create({ In this example, the `tools` array defines a `get_weather` tool, specifying its name, description, and input schema. When `tools` are provided, `expectedOutputs` must include `{ type: "tool-call" }`, and `expectedInputs` must include `{ type: "tool-response" }` so the application can send tool execution results back to the model. (`expectedInputs` only needs to include `{ type: "tool-call" }` if the application intends to pass assistant-role tool calls as input in `initialPrompts`, `append()`, or `prompt()`, such as to provide few-shot examples or restore a saved session's context.) -Few-shot examples of tool use can be appended like so: +#### Prompting with tools (Open loop) + +When a session is configured with only `"text"` in `expectedOutputs` (or when `expectedOutputs` is omitted), `session.prompt()` resolves to a `DOMString` as usual. + +When `expectedOutputs` includes non-text types such as `{ type: "tool-call" }`, `session.prompt()` consistently resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence`), even if the model only produces text for a given turn (e.g., `[{ type: "text", value: "..." }]`). If the model outputs both text and tool calls, the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. + +Each `LanguageModelToolCall` object contains: +* `callId`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callId` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. +* `name`: The name of the tool to invoke. +* `arguments`: A dictionary fitting the JSON `inputSchema` of the tool's declaration (which must have `type: "object"`). + +The application executes the requested tool and sends the result back to the session using `LanguageModelToolSuccess` (or `LanguageModelToolError` if execution failed): ```js -await session.append([ +let response = await session.prompt("What is the weather in Seattle?"); +const toolCallMsg = response.find((msg) => msg.type === "tool-call"); + +if (toolCallMsg && toolCallMsg.value.name === "get_weather") { + const toolCall = toolCallMsg.value; + const toolResult = await getWeather(toolCall.arguments.location); + + // For simplicity, this example assumes a single tool call followed by a + // final text response. In practice, the model may respond with additional + // tool calls, which would typically be handled in a loop. + response = await session.prompt([ + { + role: "user", + content: [ + { + type: "tool-response", + value: new LanguageModelToolSuccess({ + callId: toolCall.callId, + name: toolCall.name, + result: [{ type: "object", value: toolResult }], + }), + }, + ], + }, + ]); +} + +const textMsg = response.find((msg) => msg.type === "text"); +console.log(textMsg?.value); +``` + +Note that a `"tool-response"` message should immediately follow the `"tool-call"` generated by the model. + +#### Seeding or restoring tool history + +If an application wants to supply assistant-role tool calls as input—for example, to provide few-shot examples or restore a saved conversation—it must also include `{ type: "tool-call" }` in `expectedInputs`: + +```js +const sessionWithHistory = await LanguageModel.create({ + expectedInputs: [{ type: "tool-call" }, { type: "tool-response" }], + expectedOutputs: [{ type: "tool-call" }], + tools: [/* ... */], +}); + +await sessionWithHistory.append([ { role: "user", content: "What is the weather in Seattle?" }, { role: "assistant", @@ -252,7 +307,7 @@ await session.append([ ]); ``` -Note that: +Summary of tool message types and values: * Message `content` `type` supports `"tool-call"` and `"tool-response"`: * `"tool-call"` content must use `role: "assistant"` and its `value` must be a `LanguageModelToolCall` instance (`new LanguageModelToolCall({ callId, name, arguments })`). * `"tool-response"` content must use `role: "user"` and its `value` must be either a `LanguageModelToolSuccess` instance (`new LanguageModelToolSuccess({ callId, name, result })`) or a `LanguageModelToolError` instance (`new LanguageModelToolError({ callId, name, errorMessage })`). @@ -262,62 +317,11 @@ Note that: * `"audio"`: an `AudioBuffer`, `BufferSource`, or `Blob`. * `"object"`: a JSON-serializable JavaScript object. +#### Future work: Automatic tool execution (Closed loop) -#### Open Loop +> **Note:** Closed-loop (automatic) tool execution is a planned future extension and is not yet part of the specification or implemented in browsers. -Open loop is enabled by specifying `{ type: "tool-call" }` in `expectedOutputs` when the session is created. - -When a session is configured with only `"text"` in `expectedOutputs` (or when `expectedOutputs` is omitted), `session.prompt()` resolves to a `DOMString` as usual. - -When `expectedOutputs` includes non-text types such as `{ type: "tool-call" }`, `session.prompt()` consistently resolves to an array of `LanguageModelMessageContent` dictionaries (`sequence`), even if the model only produces text for a given turn (e.g., `[{ type: "text", value: "..." }]`). If the model outputs both text and tool calls, the text is included first (`{ type: "text", value: "..." }`), followed by `{ type: "tool-call", value: LanguageModelToolCall }` items. - -Each `LanguageModelToolCall` object contains: -* `callId`: An opaque string identifier for this tool call. Its format is implementation- and model-defined. Applications should not rely on any specific format and should simply pass `toolCall.callId` back in the corresponding `LanguageModelToolSuccess` or `LanguageModelToolError`. -* `name`: The name of the tool to invoke. -* `arguments`: A dictionary fitting the JSON `inputSchema` of the tool's declaration (which must have `type: "object"`). - -The client is expected to execute the tool and send the result back to the session using `LanguageModelToolSuccess` (or `LanguageModelToolError` if execution failed): - -```js -const sessionOptions = structuredClone(options); -sessionOptions.expectedInputs = [{ type: "tool-response" }]; -sessionOptions.expectedOutputs = [{ type: "tool-call" }]; -const session = await LanguageModel.create(sessionOptions); - -let result = await session.prompt("What is the weather in Seattle?"); -if (Array.isArray(result)) { - const toolCallMsg = result.find((msg) => msg.type === "tool-call"); - if (toolCallMsg && toolCallMsg.value.name === "get_weather") { - const toolCall = toolCallMsg.value; - const toolResult = getWeather(toolCall.arguments.location); - // For simplicity, this example assumes a single tool call followed by a - // final text response. In practice, the model may respond with additional - // tool calls, which would typically be handled in a loop. - result = await session.prompt([ - { - role: "user", - content: [ - { - type: "tool-response", - value: new LanguageModelToolSuccess({ - callId: toolCall.callId, - name: toolCall.name, - result: [{ type: "object", value: toolResult }], - }), - }, - ], - }, - ]); - } -} -console.log(result); -``` - -Note that a `"tool-response"` message should immediately follow the `"tool-call"` generated by the model. - -#### Closed Loop - -To enable automatic execution, add an `execute` function for each tool's implementation, and add a `toolUseConfig` to indicate that execution is enabled and pose a max number of tool calls invoked in a single session generation: +In a future closed-loop mode, tool declarations could optionally include an `execute` callback so the user agent can automatically invoke tools and feed their results back to the model within a single `session.prompt()` call: ```js const session = await LanguageModel.create({ @@ -327,8 +331,6 @@ const session = await LanguageModel.create({ content: "You are a helpful assistant. You can use tools to help the user.", }, ], - expectedInputs: [{ type: "text", languages: ["en"] }, { type: "tool-response" }], - expectedOutputs: [{ type: "text", languages: ["en"] }, { type: "tool-call" }], tools: [ { name: "get_weather", @@ -347,8 +349,7 @@ const session = await LanguageModel.create({ const res = await fetch( "https://weatherapi.example/?location=" + encodeURIComponent(location), ); - // Returns the result as a JSON string. - return JSON.stringify(await res.json()); + return await res.json(); }, }, ], @@ -357,48 +358,15 @@ const session = await LanguageModel.create({ const result = await session.prompt("What is the weather in Seattle?"); ``` -When the language model determines that a tool call is needed, the user agent invokes the `get_weather` tool's `execute()` function with the provided arguments and returns the result to the model, which can then incorporate it into its response. - -#### Do I need auto execution? - -In general, automatic execution is suitable for use cases where the model quality is good enough via prompt tuning. That can either mean you can tolerate certain mistakes that the model makes when making tool calls, or the task is simple enough for the model to handle (e.g., just a few distinct tools, short and clean tool output, short context window, etc.). - -On the other hand, open loop allows more flexibility for intercepting at various points in the planner loop (the reason -> action -> observation loop) where you can inject your business logic programmatically. - -Here are a few patterns where open loop is useful: - -1) **Context management** - -If your session might go through a long chain of interactions, and the previous tool results are no longer important or relevant for your use case, open loop gives you the flexibility of editing and recreating the session in the middle of a tool call. You can manually compress and modify the history, and recreate a new session with less content. - -For example, for a shopping agent, your tool keeps track of a live shopping cart, but only the latest cart status is important. When there have been multiple rounds of cart updates, you might need to compress the tool call history to avoid exceeding the context window and to improve latency and quality. - -2) **Conditional loop breaking** - -If your business logic requires determinism in some critical states, open loop allows the flexibility to early-exit the planner loop and output a pre-determined action. - -For example, for a shopping agent, you might be required to get an explicit confirmation before placing the order. Whenever the tool `"place_order"` is called for the first time, you can exit the planner loop immediately and display a verbatim confirmation message to the user. - -3) **Conditional constraints** - -In automatic execution, the planner loop decodes multiple times behind the scenes. If you need to supply constraints dynamically, you can use the open loop API and control the planner loop yourself, since closed loop doesn't have a natural way to supply a different `responseConstraint` for each LLM step. - -For example, you might want the model to always generate tool `FOO` after tool `BAR` is called, or you might want the model to always generate text-only output with a specific prefix after tool `FOO` is called. - -#### Concurrent tool use - -Developers should be aware that the model might call their tool multiple times, concurrently. For example, code such as: - -```js -const result = await session.prompt( - "Which of these locations currently has the highest temperature? Seattle, Tokyo, Berlin" -); -``` +##### When to use open loop vs. closed loop -might emit three `"get_weather"` tool calls in a single turn (or, in closed loop mode, call `execute()` three times concurrently and wait for all tool call results using the equivalent of `Promise.all()` internally before composing its final response). +Automatic execution (closed loop) is convenient for straightforward tasks where the application wants the user agent to run the entire reason → action → observation loop automatically until the model produces a final answer. -Similarly, the model might call multiple different tools in a single turn if it believes they are all relevant when responding to the given prompt. +Open loop is the lower-level primitive and remains necessary when the application needs fine-grained control between model turns: +1. **Context and history management:** In long-running sessions, tool outputs can quickly consume the context window. Open loop lets the application inspect `session.contextUsage`, summarize or truncate large tool results before sending them to the model, or compact stale tool call history into a fresh session (for example, keeping only the latest shopping cart state after multiple cart updates). +2. **Human-in-the-loop and non-destructive interception:** While a closed-loop `execute()` callback could abort an entire `prompt()` call via an `AbortSignal` (rejecting the promise and discarding the in-flight turn), open loop resolves with the proposed `LanguageModelToolCall` already committed to the session context. This makes it easy to pause for explicit user confirmation before executing a sensitive tool (such as `"place_order"`), and then either resume the session with `LanguageModelToolSuccess`, report a user rejection via `LanguageModelToolError` so the model can adjust, or handle the action directly in application UI without another model call. +3. **Per-step `responseConstraint` and decoding control:** Because closed loop runs multiple generation steps inside a single `prompt()` call, it cannot easily apply different decoding options to individual steps. With open loop, each step is an explicit `prompt()` call, allowing the application to pass a `responseConstraint` (JSON Schema or regex) or an assistant response `prefix` on a specific turn (for example, constraining the model's final response after a tool returns). ### Multimodal inputs From 7ce4f5bcd590c45930ef2c041d94dc8f13338745 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 13:45:48 -0700 Subject: [PATCH 27/30] add back concurrent tool use --- README.md | 12 ++++++++++++ 1 file changed, 12 insertions(+) diff --git a/README.md b/README.md index b8871e1..a249758 100644 --- a/README.md +++ b/README.md @@ -368,6 +368,18 @@ Open loop is the lower-level primitive and remains necessary when the applicatio 2. **Human-in-the-loop and non-destructive interception:** While a closed-loop `execute()` callback could abort an entire `prompt()` call via an `AbortSignal` (rejecting the promise and discarding the in-flight turn), open loop resolves with the proposed `LanguageModelToolCall` already committed to the session context. This makes it easy to pause for explicit user confirmation before executing a sensitive tool (such as `"place_order"`), and then either resume the session with `LanguageModelToolSuccess`, report a user rejection via `LanguageModelToolError` so the model can adjust, or handle the action directly in application UI without another model call. 3. **Per-step `responseConstraint` and decoding control:** Because closed loop runs multiple generation steps inside a single `prompt()` call, it cannot easily apply different decoding options to individual steps. With open loop, each step is an explicit `prompt()` call, allowing the application to pass a `responseConstraint` (JSON Schema or regex) or an assistant response `prefix` on a specific turn (for example, constraining the model's final response after a tool returns). +#### Concurrent tool use + +Developers should be aware that the model might call their tool multiple times, concurrently. For example, code such as + +```js +const result = await session.prompt("Which of these locations currently has the highest temperature? Seattle, Tokyo, Berlin"); +``` + +might call the above `"getWeather"` tool's `execute()` function three times. The model would wait for all tool call results to return, using the equivalent of `Promise.all()` internally, before it composes its final response. + +Similarly, the model might call multiple different tools, if it believes they all are relevant when responding to the given prompt. + ### Multimodal inputs All of the above examples have been of text prompts. Some language models also support other inputs. Our design initially includes the potential to support images and audio clips as inputs. This is done by using objects in the form `{ type: "image", content }` and `{ type: "audio", content }` instead of strings. The `content` values can be the following: From d5eee6cf1c8e498fe3e172dc92213e3ed1a851b2 Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 15:38:53 -0700 Subject: [PATCH 28/30] fix PR bot error --- index.bs | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/index.bs b/index.bs index fe07b83..37b1115 100644 --- a/index.bs +++ b/index.bs @@ -582,7 +582,7 @@ The following are the [=event handlers=] (and their corresponding [=event handle 1. If |hasNonTextExpectedOutput| is false, then return |promise|. - 1. Return the result of [=reacting=] to |promise| with a fulfillment handler that returns |contents|. + 1. Return the result of [=promise/reacting=] to |promise| with a fulfillment handler that returns |contents|.
@@ -1067,7 +1067,7 @@ Every {{LanguageModelToolCall}} has a call IDname, a [=string=], set during creation. -Every {{LanguageModelToolCall}} has an arguments, an [=ECMAScript/object=] or null, set during creation. +Every {{LanguageModelToolCall}} has an arguments, an [=Object=] or null, set during creation.
The new LanguageModelToolCall(|init|) constructor steps are: @@ -1091,7 +1091,7 @@ Every {{LanguageModelToolSuccess}} has a cal Every {{LanguageModelToolSuccess}} has a name, a [=string=], set during creation. -Every {{LanguageModelToolSuccess}} has a result, a [=frozen array=] of {{LanguageModelToolResultContent}}s, set during creation. +Every {{LanguageModelToolSuccess}} has a result, a {{FrozenArray}}<{{LanguageModelToolResultContent}}>, set during creation.
The new LanguageModelToolSuccess(|init|) constructor steps are: From 2b64d820155deceb5b4fe4c55c095402b7eccf7e Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 15:57:08 -0700 Subject: [PATCH 29/30] fix rendering --- index.bs | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/index.bs b/index.bs index 37b1115..673a291 100644 --- a/index.bs +++ b/index.bs @@ -288,11 +288,11 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. [=set/Append=] |tool|["{{LanguageModelToolDeclaration/name}}"] to |toolNames|. 1. If |tool|["{{LanguageModelToolDeclaration/description}}"] is the empty [=string=], then throw a {{TypeError}}. 1. Let |schema| be |tool|["{{LanguageModelToolDeclaration/inputSchema}}"]. - 1. Let |typeValue| be ? [$Get$](|schema|, "`type`"). + 1. Let |typeValue| be ? [$Get$]\(|schema|, "type"\). 1. If |typeValue| is not "`object`", then throw a {{TypeError}}. - 1. Let |propertiesValue| be ? [$Get$](|schema|, "`properties`"). + 1. Let |propertiesValue| be ? [$Get$]\(|schema|, "properties"\). 1. If |propertiesValue| is not undefined and |propertiesValue| is not an [=Object=], then throw a {{TypeError}}. - 1. Let |requiredValue| be ? [$Get$](|schema|, "`required`"). + 1. Let |requiredValue| be ? [$Get$]\(|schema|, "required"\). 1. If |requiredValue| is not undefined and ? [$IsArray$](|requiredValue|) is false, then throw a {{TypeError}}. 1. Perform ? [=serialize a JavaScript value to a JSON string=] given |schema|. From 2fdf71a2fc02259b5fc97afd3d2ee570097ab91f Mon Sep 17 00:00:00 2001 From: jingyun19 Date: Thu, 1 Oct 2026 16:01:47 -0700 Subject: [PATCH 30/30] Fix formatting issues in spec --- index.bs | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/index.bs b/index.bs index 673a291..935fb7a 100644 --- a/index.bs +++ b/index.bs @@ -288,11 +288,11 @@ typedef (LanguageModelToolSuccess or LanguageModelToolError) LanguageModelToolRe 1. [=set/Append=] |tool|["{{LanguageModelToolDeclaration/name}}"] to |toolNames|. 1. If |tool|["{{LanguageModelToolDeclaration/description}}"] is the empty [=string=], then throw a {{TypeError}}. 1. Let |schema| be |tool|["{{LanguageModelToolDeclaration/inputSchema}}"]. - 1. Let |typeValue| be ? [$Get$]\(|schema|, "type"\). + 1. Let |typeValue| be ? [$Get$]\(|schema|, "type"). 1. If |typeValue| is not "`object`", then throw a {{TypeError}}. - 1. Let |propertiesValue| be ? [$Get$]\(|schema|, "properties"\). + 1. Let |propertiesValue| be ? [$Get$]\(|schema|, "properties"). 1. If |propertiesValue| is not undefined and |propertiesValue| is not an [=Object=], then throw a {{TypeError}}. - 1. Let |requiredValue| be ? [$Get$]\(|schema|, "required"\). + 1. Let |requiredValue| be ? [$Get$]\(|schema|, "required"). 1. If |requiredValue| is not undefined and ? [$IsArray$](|requiredValue|) is false, then throw a {{TypeError}}. 1. Perform ? [=serialize a JavaScript value to a JSON string=] given |schema|.