Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
4debf5b
feat(provider): read model facts from the served catalog instead of t…
kojiwakayama Sep 27, 2026
1f2c892
fix(provider): scope served model facts per project and credential
kojiwakayama Sep 27, 2026
e3b99bd
docs(agent): describe the alias wait without the runtime kind
kojiwakayama Sep 27, 2026
b7bc757
fix(provider): resolve the served default model and keep the base URL…
kojiwakayama Sep 27, 2026
9d81678
Merge origin/main into feat/sdk-catalog-client
kojiwakayama Sep 27, 2026
43157b5
fix(provider): load the catalog before every model decision and read …
kojiwakayama Sep 27, 2026
77a262d
fix(provider): resolve served-only aliases and keep each caller's wai…
kojiwakayama Sep 27, 2026
7f5f2ff
fix(provider): refuse models only against a fresh catalog; route serv…
kojiwakayama Sep 27, 2026
6a39b6c
docs(changelog): describe served-only aliases and stale-catalog refusals
kojiwakayama Sep 27, 2026
01b6d89
Merge origin/main into feat/sdk-catalog-client
kojiwakayama Sep 27, 2026
10c0991
fix(agent): let credential-free contexts read the run's catalog by key
kojiwakayama Sep 27, 2026
a75ffb4
fix(provider): settle a served model before building or describing it
kojiwakayama Sep 27, 2026
47afa90
fix(provider): log catalog failures per scope and state the stale-cat…
kojiwakayama Sep 27, 2026
d0ee50c
fix(provider): forward members a rebuilt model gains; cancel judge ca…
kojiwakayama Sep 27, 2026
9f3ce10
fix(provider): keep the catalog cache within its cap after concurrent…
kojiwakayama Sep 27, 2026
8b14445
fix(agent): describe a served model to the executor only once it has …
kojiwakayama Sep 27, 2026
a8c9255
fix(agent): apply Anthropic defaults by served surface, not the provi…
kojiwakayama Sep 27, 2026
3c9ed06
fix(provider): never log a signed query value from a catalog failure
kojiwakayama Sep 27, 2026
383da0a
Merge remote-tracking branch 'origin/main' into feat/sdk-catalog-client
kojiwakayama Sep 27, 2026
7aa40a2
fix(runtime): record native controls by served surface
kojiwakayama Sep 27, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 41 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,47 @@ For finish-required streams in the active lifecycle, a completed turn containing
only unavailable tool calls now permits the same recovery turn as the legacy lifecycle. Rejected tools are
not executed, and malformed or empty handoff requests still fail.

### Changed: Veryfront Cloud model facts come from the served model catalog

`veryfront-cloud/*` models now read their facts from the model catalog your
Veryfront API serves at `<api>/ai/models`, instead of from a table shipped in
this package. The facts are the wire protocol, whether the Responses operation
is served, the default thinking budget, the two chat completions capability
flags, short model aliases such as `opus`, and the default model. For the
models this package listed, the served facts are the same, so requests are
unchanged.

- Model construction stays synchronous and makes no network call. The catalog
loads on the first async step of a model (`prepare`, `doGenerate` or
`doStream`), with the same credentials and project as inference, and is
cached for five minutes per API, project and credential.
- Until the catalog has loaded for the credentials in use, and whenever it
cannot be loaded, the facts shipped with this package apply, as in the
previous release. A model whose first load failed tries again on a later
call.
- A model keeps the facts it settled with for its lifetime. A catalog refreshed
later applies to models constructed after the refresh.
- Agents resolve a short alias or a provider the platform added after this
release through Veryfront Cloud once the catalog has loaded, after the
built-in aliases, so a bare vendor model name keeps its meaning for your own
provider key.
- A model the catalog does not list is refused only against a catalog loaded
within the last five minutes, or the shipped list before any has loaded. A
model enabled since the catalog was cached is not refused; the catalog is
refreshed first.
- `loadVeryfrontCloudModelCatalog()` loads the catalog for the Veryfront Cloud
credentials in effect, so synchronous helpers such as
`resolveVeryfrontCloudModelId("opus")` and
`resolveVeryfrontCloudModelThinking()` read served facts, including models
the platform added after this release.
- `resolveVeryfrontCloudDefaultModelId()` returns the default model the
catalog names, or the built-in default before it loads.
`VeryfrontCloudModelId` types a model ID as `<provider>/<model>`.
- `VERYFRONT_CLOUD_CHAT_MODELS`, `findVeryfrontCloudModel`,
`findVeryfrontCloudModelByModelId` and `groupVeryfrontCloudModelsByProvider`
are deprecated. They still return the list shipped with this package, and a
later release removes them.

### Changed: Veryfront Cloud models call the vendor-neutral endpoints

`veryfront-cloud/*` models that speak the OpenAI protocol now send requests to
Expand Down
68 changes: 36 additions & 32 deletions docs/api-reference/veryfront/provider.md

Large diffs are not rendered by default.

33 changes: 33 additions & 0 deletions src/agent/hosted/application-model-resolver.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,10 @@ import { assert, assertEquals, assertRejects, assertThrows } from "#veryfront/te
import { describe, it } from "#veryfront/testing/bdd.ts";
import { revokeModelRuntimeResolver } from "../runtime/model-transport.ts";
import { createHostedApplicationModelResolver } from "./application-model-resolver.ts";
import {
__resetVeryfrontCloudCatalogForTests,
__setVeryfrontCloudCatalogForScopeForTests,
} from "#veryfront/provider/veryfront-cloud/catalog-client.ts";

const modelId = "veryfront-cloud/openai/gpt-4o";

Expand Down Expand Up @@ -150,4 +154,33 @@ describe("hosted application model authority", () => {
"authority is revoked",
);
});
it("exposes the reconciliation hook of the protocol a model settles on", async () => {
const servedId = "veryfront-cloud/acme/acme-gemini";
const options = { ...resolverOptions(), allowedModelIds: new Set([servedId]) };
try {
const model = createHostedApplicationModelResolver(options)(servedId)!;
// Cold, an unlisted provider is built as an OpenAI-protocol model.
assertEquals(model._reconcileProviderMetadata, undefined);

__setVeryfrontCloudCatalogForScopeForTests(
{ apiBaseUrl: options.apiBaseUrl, apiToken: options.authToken },
{
models: [{
id: "acme-gemini",
modelId: "acme/acme-gemini",
provider: "acme",
surface: "google",
operations: ["chat-completions"],
aliases: [],
capabilities: {},
}],
},
);
await model.prepare!();

assert(typeof model._reconcileProviderMetadata === "function");
} finally {
__resetVeryfrontCloudCatalogForTests();
}
});
});
64 changes: 44 additions & 20 deletions src/agent/hosted/application-model-resolver.ts
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,10 @@ import {
type VeryfrontCloudContext,
} from "#veryfront/provider/veryfront-cloud/context.ts";
import { createVeryfrontCloudModel } from "#veryfront/provider/veryfront-cloud/provider.ts";
import {
readVeryfrontCloudModelFacts,
registerVeryfrontCloudModelFacts,
} from "#veryfront/provider/veryfront-cloud/model-catalog.ts";
import { requireSecureInferenceApiBaseUrl } from "#veryfront/provider/veryfront-cloud/shared.ts";
import {
type AgentModelRuntimeResolver,
Expand Down Expand Up @@ -135,15 +139,30 @@ export function createHostedApplicationModelResolver(input: {
assertCredentialActive: assertActive,
})
);
const reconcile = model._reconcileProviderMetadata;
// Metadata is read from the model on each access: a Veryfront Cloud model
// may settle a different protocol and capabilities on its first async step.
const proxy: ModelRuntime<ModelRuntimeCallOptions> = Object.freeze({
specificationVersion: model.specificationVersion,
provider: model.provider,
modelProvider: model.modelProvider,
modelId: model.modelId,
executionMode: model.executionMode,
runtimeCapabilities: model.runtimeCapabilities,
_generateViaStream: model._generateViaStream,
get specificationVersion() {
return model.specificationVersion;
},
get provider() {
return model.provider;
},
get modelProvider() {
return model.modelProvider;
},
get modelId() {
return model.modelId;
},
get executionMode() {
return model.executionMode;
},
get runtimeCapabilities() {
return model.runtimeCapabilities;
},
get _generateViaStream() {
return model._generateViaStream;
},
async prepare(abortSignal?: AbortSignal) {
await run(callScope(abortSignal), (signal) => model.prepare?.(signal));
},
Expand Down Expand Up @@ -175,19 +194,24 @@ export function createHostedApplicationModelResolver(input: {
throw error;
}
},
...(typeof reconcile === "function"
? {
async _reconcileProviderMetadata(options: {
providerMetadata: Record<string, unknown>;
suppressedToolCalls: readonly { id: string; name: string }[];
abortSignal?: AbortSignal;
}) {
return await run(callScope(options.abortSignal), (signal) =>
reconcile.call(model, { ...options, abortSignal: signal }));
},
}
: {}),
// Read when described and when invoked: a model rebuilt onto another
// protocol may gain or lose its reconciliation hook.
get _reconcileProviderMetadata() {
const reconcile = model._reconcileProviderMetadata;
if (typeof reconcile !== "function") return undefined;
return async (options: {
providerMetadata: Record<string, unknown>;
suppressedToolCalls: readonly { id: string; name: string }[];
abortSignal?: AbortSignal;
}) =>
await run(
callScope(options.abortSignal),
(signal) => reconcile.call(model, { ...options, abortSignal: signal }),
);
},
});
// The proxy reports the facts of the model it calls.
registerVeryfrontCloudModelFacts(proxy, () => readVeryfrontCloudModelFacts(model)!);
models.set(id, proxy);
return proxy;
};
Expand Down
32 changes: 24 additions & 8 deletions src/agent/hosted/cloud-chat-execution-preparation.ts
Original file line number Diff line number Diff line change
@@ -1,5 +1,12 @@
import { resolveVeryfrontCloudModelThinking } from "#veryfront/provider";
import { runWithVeryfrontCloudContext } from "#veryfront/provider/veryfront-cloud/context.ts";
import {
runWithVeryfrontCloudContext,
runWithVeryfrontCloudContextAsync,
} from "#veryfront/provider/veryfront-cloud/context.ts";
import { loadVeryfrontCloudModelCatalog } from "#veryfront/provider/veryfront-cloud/shared.ts";

/** Longest preparation waits for the served catalog before resolving the model. */
const CATALOG_MAX_WAIT_MS = 3_000;
import { resolveRuntimeModel } from "../runtime/model-resolution.ts";
import type { HostedChatRuntimeCreationResult } from "./chat-runtime-contract.ts";
import {
Expand Down Expand Up @@ -88,16 +95,25 @@ export async function prepareVeryfrontCloudHostedChatExecution<
HostedChatExecutionPreparationResult<TRuntimeAgentDefinition, TRuntimeResult>
> {
const { logger, rootRun, ...preparationInput } = input;
const cloudContext = {
apiBaseUrl: String(input.apiUrl),
apiToken: input.request.authToken,
projectSlug: input.request.projectSlug,
serviceLayer: "cloud" as const,
};
// Model ids and thinking defaults resolve against the served catalog, loaded
// and read under the request's own credentials and project.
await runWithVeryfrontCloudContextAsync(
cloudContext,
() => loadVeryfrontCloudModelCatalog({ maxWaitMs: CATALOG_MAX_WAIT_MS }),
);
const resolveModelId = (modelId: string | undefined): string | undefined =>
runWithVeryfrontCloudContext(
{
apiBaseUrl: String(input.apiUrl),
apiToken: input.request.authToken,
projectSlug: input.request.projectSlug,
serviceLayer: "cloud",
},
cloudContext,
() => modelId === undefined ? undefined : resolveRuntimeModel(modelId),
);
const resolveModelThinking: typeof resolveVeryfrontCloudModelThinking = (modelId) =>
runWithVeryfrontCloudContext(cloudContext, () => resolveVeryfrontCloudModelThinking(modelId));

return await prepareHostedChatExecution({
...preparationInput,
Expand All @@ -106,6 +122,6 @@ export async function prepareVeryfrontCloudHostedChatExecution<
logger,
}),
resolveModelId,
resolveModelThinking: resolveVeryfrontCloudModelThinking,
resolveModelThinking,
});
}
40 changes: 32 additions & 8 deletions src/agent/hosted/context-summary-generator.ts
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,12 @@ import {
resolveVeryfrontCloudGatewayModelId,
resolveVeryfrontCloudModelId,
} from "../../provider/index.ts";
import { runWithVeryfrontCloudContextAsync } from "#veryfront/provider/veryfront-cloud/context.ts";
import {
runWithVeryfrontCloudContext,
runWithVeryfrontCloudContextAsync,
type VeryfrontCloudContext,
} from "#veryfront/provider/veryfront-cloud/context.ts";
import { loadVeryfrontCloudModelCatalog } from "#veryfront/provider/veryfront-cloud/shared.ts";
import { generateText } from "../../runtime/runtime-bridge.ts";
import { redactSensitive, sanitizeUrlCredentials } from "#veryfront/utils";
import type { TextGenerationRuntimeMessage } from "../runtime/text-generation-runtime-message-types.ts";
Expand Down Expand Up @@ -168,12 +173,7 @@ async function summarizeSegment(input: {
const generate = input.options.generateText ?? generateText;
const resolve = input.options.resolveModel ?? resolveModel;
const result = await runWithVeryfrontCloudContextAsync(
{
apiBaseUrl: input.options.apiUrl.toString(),
apiToken: input.options.authToken,
projectSlug: input.options.projectSlug ?? undefined,
serviceLayer: "cloud",
},
summaryCloudContext(input.options),
() =>
Promise.resolve(generate({
model: resolve(input.modelId),
Expand All @@ -192,6 +192,20 @@ async function summarizeSegment(input: {
return result.text.trim();
}

function summaryCloudContext(
options: VeryfrontCloudContextSummaryGeneratorOptions,
): VeryfrontCloudContext {
return {
apiBaseUrl: options.apiUrl.toString(),
apiToken: options.authToken,
projectSlug: options.projectSlug ?? undefined,
serviceLayer: "cloud",
};
}

/** Longest summary generation waits for the served catalog before resolving its model. */
const CATALOG_MAX_WAIT_MS = 3_000;

function resolveSummaryModelId(model: string | undefined): string {
const cloudModelId = resolveVeryfrontCloudModelId(model);
return resolveVeryfrontCloudGatewayModelId(cloudModelId) ?? cloudModelId;
Expand All @@ -202,7 +216,17 @@ export function createVeryfrontCloudContextSummaryGenerator(
options: VeryfrontCloudContextSummaryGeneratorOptions,
): ContextSummaryGenerator {
return async ({ messagesToSummarize, retainedMessages, customInstructions }) => {
const modelId = resolveSummaryModelId(options.model);
// The model resolves against the served catalog loaded for the same
// credentials and project the summary calls use.
const cloudContext = summaryCloudContext(options);
await runWithVeryfrontCloudContextAsync(
cloudContext,
() => loadVeryfrontCloudModelCatalog({ maxWaitMs: CATALOG_MAX_WAIT_MS }),
);
Comment thread
kojiwakayama marked this conversation as resolved.
const modelId = runWithVeryfrontCloudContext(
cloudContext,
() => resolveSummaryModelId(options.model),
);
const chunks = chunkSerializedMessages(messagesToSummarize, options.maxInputTokens);
let summary = "";

Expand Down
4 changes: 4 additions & 0 deletions src/agent/hosted/default-chat-runtime.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,7 @@ import {
assertStringIncludes,
} from "#veryfront/testing/assert.ts";
import { it } from "#veryfront/testing/bdd.ts";
import { useServedCatalogForTests } from "#veryfront/provider/veryfront-cloud/catalog-client.test-helpers.ts";
import { deleteEnv, getEnv, setEnv } from "#veryfront/compat/process.ts";
import { refreshEnvironmentConfig } from "#veryfront/config/environment-config.ts";
import { clearModelProviders, type ModelRuntime, registerModelProvider } from "#veryfront/provider";
Expand Down Expand Up @@ -465,6 +466,7 @@ it("applies refreshed structured system messages in hosted chat", async () => {
});

Deno.test("createDefaultHostedChatRuntime builds a cloud-backed hosted runtime", async () => {
using _catalog = useServedCatalogForTests();
let capturedContext: DefaultHostedChatRuntimeTaskContext | undefined;
let capturedCapability: unknown;
const runEventWriterCapability = createHostedRunEventWriterCapability({
Expand Down Expand Up @@ -1027,6 +1029,7 @@ Deno.test("hosted first provider call filters skill tools for every tool selecto
});

Deno.test("createDefaultHostedChatRuntime forwards hosted project slug to integration discovery", async () => {
using _catalog = useServedCatalogForTests();
const previousApiBaseUrl = getEnv("VERYFRONT_API_BASE_URL");
const previousApiToken = getEnv("VERYFRONT_API_TOKEN");
const previousProjectSlug = getEnv("VERYFRONT_PROJECT_SLUG");
Expand Down Expand Up @@ -1106,6 +1109,7 @@ Deno.test("createDefaultHostedChatRuntime forwards hosted project slug to integr
});

Deno.test("createDefaultHostedChatRuntime keeps per-run host tools out of the global registry", async () => {
using _catalog = useServedCatalogForTests();
try {
const createRuntime = (description: string) =>
createDefaultHostedChatRuntime({
Expand Down
Loading
Loading