{"title":"GenAI Management Proxy — LLM Provider Selection and Integration Guide","version":"1.0.0","lastUpdated":"2026-10-03","baseUrl":"https://oauth.xiangenhu.info","domainPolicy":{"mode":"fixed","origin":"https://oauth.xiangenhu.info","deriveFromRequest":false,"instructions":"Use this literal gateway origin in consuming apps. Do not infer it from Host, X-Forwarded-Host, window.location or the consuming app domain.","deploymentRequirement":"Set GENAI_PUBLIC_ORIGIN=https://oauth.xiangenhu.info on the gateway. If existing accounts use another origin, plan a migration before changing it."},"links":{"json":"https://oauth.xiangenhu.info/genaiInstruction","jsonFile":"https://oauth.xiangenhu.info/genaiInstruction.json","html":"https://oauth.xiangenhu.info/genaiInstruction.html","text":"https://oauth.xiangenhu.info/genaiInstruction.txt","sdk":"https://oauth.xiangenhu.info/genaiInstruction/sdk.js","portal":"https://oauth.xiangenhu.info/genai/","oauthAndSMTP":"https://oauth.xiangenhu.info/instruction"},"preserve":["Existing OAuth login","Existing user identities and permissions","Existing SMTP integration"],"providerSelection":{"location":"User profile → AI models and providers","choices":[{"id":"hosted","label":"Built-in model","billing":"Configured allowance or credit; requires operator setup"},{"id":"personal","label":"My provider and API key","billing":"Provider bills the user directly"},{"id":"local","label":"My local LLM","requirement":"Authenticated public HTTPS endpoint reachable by the gateway"}],"catalog":[{"id":"openai","label":"OpenAI","baseURL":"https://api.openai.com/v1","docsURL":"https://platform.openai.com/docs/api-reference","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.openai.com/v1"}]},{"id":"anthropic","label":"Anthropic / Claude","baseURL":"https://api.anthropic.com/v1","docsURL":"https://docs.anthropic.com/en/api/models-list","category":"Direct APIs","protocol":"Anthropic Messages","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.anthropic.com/v1"}]},{"id":"azure","label":"Microsoft Azure OpenAI","docsURL":"https://learn.microsoft.com/azure/ai-services/openai/","category":"Direct APIs","protocol":"Azure OpenAI","access":"direct","discovery":false,"notes":"Use the resource endpoint and your deployment name, not a public model name.","endpoints":[]},{"id":"google","label":"Google Gemini","baseURL":"https://generativelanguage.googleapis.com/v1beta/openai","docsURL":"https://ai.google.dev/gemini-api/docs/openai","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://generativelanguage.googleapis.com/v1beta/openai"}]},{"id":"deepseek","label":"DeepSeek","baseURL":"https://api.deepseek.com","docsURL":"https://api-docs.deepseek.com/","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.deepseek.com"}]},{"id":"zai","label":"Z.ai / Zhipu (智谱) / GLM","baseURL":"https://api.z.ai/api/paas/v4","docsURL":"https://docs.z.ai/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":false,"endpoints":[{"label":"Z.ai international","url":"https://api.z.ai/api/paas/v4"},{"label":"Zhipu / BigModel China mainland","url":"https://open.bigmodel.cn/api/paas/v4"}]},{"id":"xai","label":"xAI","baseURL":"https://api.x.ai/v1","docsURL":"https://docs.x.ai/","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["grok-4","grok-4-fast-non-reasoning"],"endpoints":[{"label":"Default API","url":"https://api.x.ai/v1"}]},{"id":"mistral","label":"Mistral","baseURL":"https://api.mistral.ai/v1","docsURL":"https://docs.mistral.ai/","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["mistral-large-latest","mistral-small-latest"],"endpoints":[{"label":"Default API","url":"https://api.mistral.ai/v1"}]},{"id":"cohere","label":"Cohere","baseURL":"https://api.cohere.ai/compatibility/v1","docsURL":"https://docs.cohere.com/docs/compatibility-api","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["command-a-03-2025","command-r7b-12-2024"],"endpoints":[{"label":"Default API","url":"https://api.cohere.ai/compatibility/v1"}]},{"id":"groq","label":"Groq","baseURL":"https://api.groq.com/openai/v1","docsURL":"https://console.groq.com/docs/openai","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["llama-3.3-70b-versatile","openai/gpt-oss-120b"],"endpoints":[{"label":"Default API","url":"https://api.groq.com/openai/v1"}]},{"id":"together","label":"Together AI","baseURL":"https://api.together.xyz/v1","docsURL":"https://docs.together.ai/","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["meta-llama/Llama-3.3-70B-Instruct-Turbo"],"endpoints":[{"label":"Default API","url":"https://api.together.xyz/v1"}]},{"id":"fireworks","label":"Fireworks AI","baseURL":"https://api.fireworks.ai/inference/v1","docsURL":"https://docs.fireworks.ai/","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["accounts/fireworks/models/llama-v3p3-70b-instruct"],"endpoints":[{"label":"Default API","url":"https://api.fireworks.ai/inference/v1"}]},{"id":"perplexity","label":"Perplexity","baseURL":"https://api.perplexity.ai","docsURL":"https://docs.perplexity.ai/","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["sonar-pro","sonar"],"endpoints":[{"label":"Default API","url":"https://api.perplexity.ai"}]},{"id":"openrouter","label":"OpenRouter","baseURL":"https://openrouter.ai/api/v1","docsURL":"https://openrouter.ai/docs/quickstart","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["anthropic/claude-sonnet-4.5","openai/gpt-4.1-mini"],"notes":"The selected aggregator may route to downstream providers according to your account policy. Review its privacy and routing settings.","endpoints":[{"label":"Default API","url":"https://openrouter.ai/api/v1"}]},{"id":"moonshot","label":"Moonshot (Kimi)","baseURL":"https://api.moonshot.ai/v1","docsURL":"https://platform.moonshot.ai/docs/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["kimi-k2-0905-preview","kimi-k2-turbo-preview"],"endpoints":[{"label":"International","url":"https://api.moonshot.ai/v1"},{"label":"China mainland","url":"https://api.moonshot.cn/v1"}]},{"id":"dashscope","label":"Alibaba DashScope (Qwen)","baseURL":"https://dashscope-intl.aliyuncs.com/compatible-mode/v1","docsURL":"https://www.alibabacloud.com/help/en/model-studio/base-url","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["qwen-plus","qwen-max","qwen-turbo"],"endpoints":[{"label":"International","url":"https://dashscope-intl.aliyuncs.com/compatible-mode/v1"},{"label":"China mainland","url":"https://dashscope.aliyuncs.com/compatible-mode/v1"},{"label":"US","url":"https://dashscope-us.aliyuncs.com/compatible-mode/v1"},{"label":"Hong Kong","url":"https://cn-hongkong.dashscope.aliyuncs.com/compatible-mode/v1"}]},{"id":"ark","label":"Volcengine Ark (Doubao)","baseURL":"https://ark.cn-beijing.volces.com/api/v3","docsURL":"https://www.volcengine.com/docs/82379","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["doubao-seed-1-6-250615","doubao-seed-1-6-flash-250615"],"endpoints":[{"label":"Default API","url":"https://ark.cn-beijing.volces.com/api/v3"}]},{"id":"qianfan","label":"Baidu Qianfan (ERNIE)","baseURL":"https://qianfan.baidubce.com/v2","docsURL":"https://cloud.baidu.com/doc/qianfan/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["ernie-4.5-turbo-128k"],"endpoints":[{"label":"Default API","url":"https://qianfan.baidubce.com/v2"}]},{"id":"hunyuan","label":"Tencent Hunyuan","baseURL":"https://api.hunyuan.cloud.tencent.com/v1","docsURL":"https://cloud.tencent.com/document/product/1729","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["hunyuan-turbos-latest","hunyuan-lite"],"endpoints":[{"label":"Default API","url":"https://api.hunyuan.cloud.tencent.com/v1"}]},{"id":"spark","label":"iFlytek Spark","baseURL":"https://spark-api-open.xf-yun.com/v1","docsURL":"https://www.xfyun.cn/doc/spark/Web.html","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["4.0Ultra","generalv3.5","lite"],"notes":"Use the provider’s APIKey:APISecret bearer credential.","endpoints":[{"label":"Default API","url":"https://spark-api-open.xf-yun.com/v1"}]},{"id":"minimax","label":"MiniMax","baseURL":"https://api.minimax.io/v1","docsURL":"https://platform.minimax.io/docs/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["MiniMax-M2","MiniMax-M1"],"endpoints":[{"label":"International","url":"https://api.minimax.io/v1"},{"label":"China mainland","url":"https://api.minimaxi.com/v1"}]},{"id":"stepfun","label":"StepFun","baseURL":"https://api.stepfun.com/v1","docsURL":"https://platform.stepfun.com/docs/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["step-2-16k","step-2-mini"],"endpoints":[{"label":"Default API","url":"https://api.stepfun.com/v1"}]},{"id":"baichuan","label":"Baichuan","baseURL":"https://api.baichuan-ai.com/v1","docsURL":"https://platform.baichuan-ai.com/docs","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["Baichuan4-Turbo","Baichuan4-Air"],"endpoints":[{"label":"Default API","url":"https://api.baichuan-ai.com/v1"}]},{"id":"yi","label":"01.AI (Yi)","baseURL":"https://api.lingyiwanwu.com/v1","docsURL":"https://platform.lingyiwanwu.com/docs","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["yi-lightning"],"endpoints":[{"label":"Default API","url":"https://api.lingyiwanwu.com/v1"}]},{"id":"sensenova","label":"SenseTime SenseNova","baseURL":"https://api.sensenova.cn/compatible-mode/v1","docsURL":"https://platform.sensenova.cn/doc","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["SenseChat-5","SenseChat-Turbo"],"endpoints":[{"label":"Default API","url":"https://api.sensenova.cn/compatible-mode/v1"}]},{"id":"siliconflow","label":"SiliconFlow","baseURL":"https://api.siliconflow.cn/v1","docsURL":"https://docs.siliconflow.com/","category":"Regional APIs","protocol":"OpenAI-compatible","access":"direct","discovery":true,"suggestions":["Qwen/Qwen3-235B-A22B-Instruct-2507","deepseek-ai/DeepSeek-V3","moonshotai/Kimi-K2-Instruct"],"endpoints":[{"label":"China mainland","url":"https://api.siliconflow.cn/v1"},{"label":"International","url":"https://api.siliconflow.com/v1"}]},{"id":"cerebras","label":"Cerebras","baseURL":"https://api.cerebras.ai/v1","docsURL":"https://inference-docs.cerebras.ai/resources/openai","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.cerebras.ai/v1"}]},{"id":"deepinfra","label":"DeepInfra","baseURL":"https://api.deepinfra.com/v1/openai","docsURL":"https://docs.deepinfra.com/chat/overview","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.deepinfra.com/v1/openai"}]},{"id":"sambanova","label":"SambaNova","baseURL":"https://api.sambanova.ai/v1","docsURL":"https://docs.sambanova.ai/docs/en/get-started/quickstart","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://api.sambanova.ai/v1"}]},{"id":"nvidia","label":"NVIDIA NIM / API Catalog","baseURL":"https://integrate.api.nvidia.com/v1","docsURL":"https://docs.api.nvidia.com/nim/reference/llm-apis","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"endpoints":[{"label":"Default API","url":"https://integrate.api.nvidia.com/v1"}]},{"id":"huggingface","label":"Hugging Face Inference Providers","baseURL":"https://router.huggingface.co/v1","docsURL":"https://huggingface.co/docs/inference-providers/en/tasks/chat-completion","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":true,"notes":"Use a token with Inference Providers permission. Downstream provider routing follows your Hugging Face account configuration.","endpoints":[{"label":"Default API","url":"https://router.huggingface.co/v1"}]},{"id":"ai21","label":"AI21 Labs / Jamba","baseURL":"https://api.ai21.com/studio/v1","docsURL":"https://docs.ai21.com/docs/jamba-foundation-models","category":"Direct APIs","protocol":"OpenAI-compatible","access":"direct","discovery":false,"endpoints":[{"label":"Default API","url":"https://api.ai21.com/studio/v1"}]},{"id":"cloudflare","label":"Cloudflare Workers AI","docsURL":"https://developers.cloudflare.com/workers-ai/configuration/open-ai-compatibility/","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":false,"customEndpoint":true,"notes":"Use https://api.cloudflare.com/client/v4/accounts/YOUR_ACCOUNT_ID/ai/v1 and a Workers AI API token.","endpoints":[]},{"id":"bedrock","label":"Amazon Bedrock","docsURL":"https://docs.aws.amazon.com/bedrock/","category":"Enterprise gateways","protocol":"OpenAI-compatible","access":"gateway","discovery":true,"customEndpoint":true,"notes":"Requires your institution’s OpenAI-compatible HTTPS gateway and gateway bearer token. Native cloud IAM, OAuth, SigV4 and service-account credentials are not accepted here.","endpoints":[]},{"id":"vertex","label":"Google Vertex AI","docsURL":"https://cloud.google.com/vertex-ai/generative-ai/docs","category":"Enterprise gateways","protocol":"OpenAI-compatible","access":"gateway","discovery":true,"customEndpoint":true,"notes":"Requires your institution’s OpenAI-compatible HTTPS gateway and gateway bearer token. Native cloud IAM, OAuth, SigV4 and service-account credentials are not accepted here.","endpoints":[]},{"id":"watsonx","label":"IBM watsonx.ai","docsURL":"https://www.ibm.com/docs/en/watsonx-as-a-service","category":"Enterprise gateways","protocol":"OpenAI-compatible","access":"gateway","discovery":true,"customEndpoint":true,"notes":"Requires your institution’s OpenAI-compatible HTTPS gateway and gateway bearer token. Native cloud IAM, OAuth, SigV4 and service-account credentials are not accepted here.","endpoints":[]},{"id":"sagemaker","label":"Amazon SageMaker","docsURL":"https://docs.aws.amazon.com/sagemaker/","category":"Enterprise gateways","protocol":"OpenAI-compatible","access":"gateway","discovery":true,"customEndpoint":true,"notes":"Requires your institution’s OpenAI-compatible HTTPS gateway and gateway bearer token. Native cloud IAM, OAuth, SigV4 and service-account credentials are not accepted here.","endpoints":[]},{"id":"foundry","label":"Microsoft Foundry / Azure AI","docsURL":"https://learn.microsoft.com/azure/ai-foundry/","category":"Enterprise gateways","protocol":"OpenAI-compatible","access":"gateway","discovery":true,"customEndpoint":true,"notes":"Requires your institution’s OpenAI-compatible HTTPS gateway and gateway bearer token. Native cloud IAM, OAuth, SigV4 and service-account credentials are not accepted here.","endpoints":[]},{"id":"databricks","label":"Databricks Model Serving","docsURL":"https://docs.databricks.com/en/machine-learning/model-serving/","category":"Inference platforms","protocol":"OpenAI-compatible","access":"direct","discovery":false,"customEndpoint":true,"notes":"Use the OpenAI-compatible serving base URL for your workspace and an authorized bearer token.","endpoints":[]},{"id":"custom","label":"Other commercial vendor / custom gateway","category":"Custom endpoints","protocol":"OpenAI-compatible","access":"direct","discovery":true,"customEndpoint":true,"notes":"For an unlisted vendor, enter its OpenAI-compatible API base URL. This supports chat-completions bearer-token APIs, not every proprietary API protocol.","endpoints":[]},{"id":"local","label":"Local LLM: Ollama, LM Studio, vLLM, llama.cpp","docsURL":"https://docs.ollama.com/api/openai-compatibility","category":"Local models","protocol":"OpenAI-compatible","access":"direct","discovery":true,"customEndpoint":true,"notes":"Use an authenticated public HTTPS tunnel to your runtime. The cloud proxy cannot connect to localhost or LAN addresses. This is local compute, not an offline connection.","endpoints":[]}],"catalogCaveat":"Templates do not guarantee model access. Enterprise gateway templates are not native IAM integrations.","wizardSteps":["Choose funding","Choose provider and region","Configure endpoint and credentials in the gateway","Discover models or enter a model/deployment ID","Review and save","Test connection","Authorize the consuming app"],"savedConnectionActions":["Select","Edit","Test connection","Remove"],"selectionRules":["Only use connections authorized for the app","Omit connection to use the app grant default","Account default changes do not change existing grants; relink and revoke the old grant","Never silently fall back to host credentials"],"testNotice":"Tests make a real provider request and may incur charges."},"endpoints":[{"method":"GET","url":"https://oauth.xiangenhu.info/genai/connect","purpose":"User consent with state and PKCE S256"},{"method":"POST","url":"https://oauth.xiangenhu.info/genai/oauth/token","authentication":"Basic app credentials","purpose":"Code exchange or rotating refresh"},{"method":"GET","url":"https://oauth.xiangenhu.info/genai/v1/connections","authentication":"GenAI user-grant Bearer token"},{"method":"POST","url":"https://oauth.xiangenhu.info/genai/v1/chat/completions","authentication":"GenAI user-grant Bearer token"},{"method":"GET","url":"https://oauth.xiangenhu.info/genai/v1/operations/:requestId","authentication":"GenAI user-grant Bearer token"}],"sections":[{"id":"overview","title":"1. Keep your existing login","text":["GenAI is an optional addition to the existing OAuth and SMTP gateway. Keep your app’s current login, user sessions, roles, email relay and OAuth callbacks. Add a “Connect GenAI Management” action in profile settings; linking AI access is separate from signing in.","Your app backend sends text requests to the gateway. The gateway uses a connection explicitly approved by the user. Provider API keys stay encrypted in the gateway; they are never returned to your app.","This guide describes the implemented v1 contract, not a generic OpenAI drop-in API. The service accepts messages, connection, maxTokens, requestId and stream. Its response and SSE events use the shapes below."]},{"id":"origin","title":"2. Use the fixed gateway origin","text":["Fixed GenAI origin: https://oauth.xiangenhu.info. Use this origin consistently for linking, management, token exchange and inference.","oauth.xiangenhu.info and oauth.skoonline.org may serve the same deployment, and both can expose this documentation. Browser cookies and sessionStorage are still origin-specific. Do not begin linking on one domain and finish on the other. Always use https://oauth.xiangenhu.info; never derive the gateway domain from Host, forwarded headers, window.location or the consuming app domain.","If a browser reports a missing login state, check GENAI_PUBLIC_ORIGIN and the domain used to start the portal. Changing the configured origin also changes the current GenAI account namespace; it is a migration, not just an alias switch."]},{"id":"register","title":"3. Register the consuming application","text":["A gateway administrator opens https://oauth.xiangenhu.info/genai/, then Administration → Register an application. Choose a stable app ID (for example vhs), a display name, and exact HTTPS callback URLs owned by that app.","Save the one-time app secret in your app backend’s secret configuration. Callback URLs must have no query string or fragment. Register production and staging explicitly. Never put the app secret in frontend JavaScript, localStorage, a URL, or a Git commit.","There are two callback settings: the consuming app callback is registered in GenAI Administration. The gateway’s own /genai/auth/callback is used by the existing OAuth redirect flow. If your deployment enforces an OAuth redirect allow-list, allow that callback there. This branch does not enforce REDIRECT_URI_ALLOWLIST itself."],"code":"GENAI_ORIGIN=https://oauth.xiangenhu.info\nGENAI_CLIENT_ID=your-registered-app-id\nGENAI_CLIENT_SECRET=<secret-injected-on-your-server>\nGENAI_CALLBACK_URL=https://yourapp.example/auth/genai/callback"},{"id":"user","title":"4. Let the user prepare and test a connection","text":["In the GenAI portal, the user selects a provider/region, enters a model or Azure deployment name, saves their API key, and clicks Test connection. Existing connections can be edited, reused and removed. Tests make a small real provider request and can incur provider charges.","Personal provider: the vendor bills the user’s provider account directly. Local model: use an authenticated public HTTPS endpoint to Ollama, LM Studio, vLLM or another compatible runtime. A cloud proxy cannot reach the user’s localhost or private LAN.","Hosted model: available only when the operator configures it. Shared allowance or credit funds requests. This version has manual verified-credit administration, not online checkout. Administrators follow the same funding rules.","The 41 templates include direct APIs and compatible gateways. Bedrock, Vertex, Foundry, SageMaker and watsonx templates require a compatible gateway; they are not native IAM adapters. A catalog entry is not proof of current model access. Discover IDs from a saved connection or enter them manually, then test."]},{"id":"sdk","title":"5. Install the server-side client","text":["Download the SDK using the link at the top of this page and save it as genai-client.js in the consuming app backend. It is a CommonJS module for Node.js 22 or later and uses built-in fetch and crypto. No npm package publication is required.","Pass only the origin, without /genai; the SDK adds the prefix. Keep the downloaded version pinned with your app and review updates before replacing it."],"code":"const { GenAIClient } = require('./genai-client');\nconst client = new GenAIClient({\n  origin: 'https://oauth.xiangenhu.info',\n  clientId: process.env.GENAI_CLIENT_ID,\n  clientSecret: process.env.GENAI_CLIENT_SECRET,\n});"},{"id":"link","title":"6. Link through explicit user consent","text":["Run linking only from an already-authenticated app session. Protect the start action using the app’s existing CSRF mechanism. Generate the URL, state and PKCE verifier with client.link(callbackURL). Persist the transaction server-side under the current app user and session before redirecting.","The portal asks which connections the app may use and which is its default. Logging in alone never authorizes API-key use or spending. Existing gateway tokens and suite identity tokens are not GenAI inference tokens.","At the app callback, require the same authenticated app session; check the saved owner, expiry and returned state, and atomically consume the transaction once. Reject mismatches before exchanging the code. Do not trust a callback email or user ID. Use a dedicated GenAI callback; leave the existing OAuth callback unchanged.","After exchange, encrypt and persist the returned access/refresh tokens under that app user. Never expose them in frontend code or logs. The transaction store and token store below are integration adapters your app must supply; the fragments are not a complete session or persistence implementation."],"code":"// Start: after existing app authentication and CSRF checks.\nconst link = client.link(process.env.GENAI_CALLBACK_URL);\nawait linkStore.save({\n  owner: currentUser.id, sessionId: currentSession.id,\n  state: link.state, verifier: link.verifier,\n  expiresAt: Date.now() + 10 * 60 * 1000,\n});\nres.redirect(link.url);\n\n// Callback: consumeForSession must verify owner, session, state, expiry\n// and atomically reject replay. It must throw on any mismatch.\nconst pending = await linkStore.consumeForSession({\n  owner: currentUser.id, sessionId: currentSession.id,\n  returnedState: req.query.state,\n});\nconst tokens = await client.exchange({\n  code: req.query.code,\n  redirectUri: process.env.GENAI_CALLBACK_URL,\n  verifier: pending.verifier,\n});\nawait tokenStore.saveEncrypted(currentUser.id, tokens);\nres.redirect('/profile');"},{"id":"inference","title":"7. Send an authorized request","text":["Obtain the user’s server-stored access token and list authorized connections with client.connections(accessToken). Omit connection to use the app grant’s default, or send an approved connection ID. Changing the account-wide default does not change an existing app grant. Relink to change the app’s permissions/default, then revoke its old grant.","Generate and durably persist the request ID before sending. The ID starts with the current UTC YYYY-MM followed by underscore and a unique suffix. Keep it with your app’s operation record for diagnosis.","Messages support system, user and assistant roles with text content: 1–100 messages, at most 100,000 characters in their serialized array. maxTokens must be an integer from 1 to 8192. Images, audio, tool calls, embeddings and arbitrary vendor parameters are not exposed by this version.","The gateway does not persist prompts or generated text in its usage ledger. Your app decides whether to store the returned response under its own privacy and retention policy."],"code":"const requestId = client.requestId();\nawait operationStore.create({ owner: currentUser.id, requestId });\nconst result = await client.complete(accessToken, {\n  requestId,\n  messages: [{ role: 'user', content: 'Explain this experiment.' }],\n  maxTokens: 1024,\n  // connection: approvedConnectionId, // optional\n});\n// result: {content, usage: {input, output} | null, requestId,\n//          connection, model, state: 'completed' | 'uncertain'}\nawait operationStore.recordResult(requestId, result);"},{"id":"stream","title":"8. Stream text and handle incomplete outcomes","text":["client.stream sends stream:true and parses named SSE events. delta carries {text}; done carries usage, requestId, connection, model and state. error carries a safe code/message. Forward text to your frontend using your app’s existing transport.","A stream ending without done is not proof that inference failed or was free. Missing provider usage produces state:uncertain. Inspect the operation before deciding what to do. A disconnected client may still leave a provider request running."],"code":"await client.stream(accessToken, {\n  requestId, // already persisted; do not reuse an ID from another operation\n  messages: [{ role: 'user', content: 'Explain this experiment.' }],\n  maxTokens: 1024,\n}, async (event, data) => {\n  if (event === 'delta') sendTextToBrowser(data.text);\n  if (event === 'done') await operationStore.recordMetadata(requestId, data);\n  // The SDK throws on error events or a stream without a final done event.\n});\n\n// After a network interruption or duplicate-request response:\nconst status = await client.operation(accessToken, requestId);"},{"id":"tokens","title":"9. Refresh, revoke and avoid duplicate spending","text":["Access tokens last 15 minutes. Refresh tokens rotate on use and expire no later than 30 days or the grant expiry. Serialize refreshes per app user across backend instances and atomically replace stored tokens. A lost exchange/refresh response can require relinking.","Authorization codes are one-use, app-bound, redirect-bound and PKCE-bound, with a two-minute expiry. App grants last 30 days. Users can revoke access in Applications; admins can disable an app. The next request checks the latest grant/app state. Revocation does not cancel already-running requests.","A used request ID returns HTTP 409; it does not replay the response. Do not silently retry inference with a new ID or switch to a host key. Check operation status, explain uncertainty, and require an explicit user decision if another generation is needed.","Previous-month request IDs are rejected for new inference. Older settled operation details can be compacted by the current ledger; keep required audit records in the consuming app. Pending and uncertain reservations remain held until verified reconciliation."],"code":"// Inside a distributed per-user refresh lock:\nconst nextTokens = await client.refresh(savedTokens.refresh_token);\nawait tokenStore.replaceEncryptedAtomically(currentUser.id, nextTokens);"},{"id":"api","title":"10. HTTP endpoint reference","text":["All paths below use the configured origin. Browser management APIs use the GenAI session plus same-origin protections. App inference APIs use GenAI user-grant Bearer tokens, never gateway login JWTs.","Basic app credentials are base64(clientId + \":\" + clientSecret). They identify the app at token exchange; credentials alone cannot select an arbitrary user or access personal keys."],"code":"GET  /genai/                          Portal\nGET  /genai/connect                   Consent (client_id, redirect_uri, state,\n                                     code_challenge, code_challenge_method=S256)\nPOST /genai/oauth/token               Basic app credentials; JSON body\nGET  /genai/v1/connections            Bearer user-grant access token\nPOST /genai/v1/chat/completions       Bearer user-grant access token; JSON body\nGET  /genai/v1/operations/:requestId   Bearer user-grant access token\n\n// Authorization-code exchange JSON:\n{ \"grant_type\": \"authorization_code\", \"code\": \"<code>\",\n  \"redirect_uri\": \"https://yourapp.example/auth/genai/callback\",\n  \"code_verifier\": \"<server-stored-verifier>\" }\n\n// Refresh JSON:\n{ \"grant_type\": \"refresh_token\", \"refresh_token\": \"<stored-refresh-token>\" }"},{"id":"errors","title":"11. Diagnostics and recovery","text":["401 access_expired: refresh once through the serialized backend refresh path. Invalid refresh or revoked grant: reconnect. 403 grant_scope: choose an approved connection or relink.","402 credit_required: the hosted allowance/credit cannot fund the reservation. Offer a personal provider or administrator-assisted credit. Do not bypass funding for admins.","409 duplicate_request or request_id_conflict: inspect the original operation. 409 choose_connection or hosted_unavailable: return the user to connection settings.","429 pending_limit or monthly_request_limit: inspect unresolved operations or the initial-release capacity limits. 503 inference_busy: no provider work was started by that rejected request; communicate capacity to the user.","provider_auth, provider_model, provider_limit and provider_unavailable describe upstream key/access, endpoint/model, quota/rate, or network/provider failures. Use portal Test connection. Do not display raw vendor error bodies or credentials.","The portal may show genai_setup or an enablement notice until the administrator configures the module. Availability of this public instruction page does not prove inference is enabled."]},{"id":"deployment","title":"12. Operator setup and release acceptance","text":["Enable GENAI_ENABLED=true; set GENAI_PUBLIC_ORIGIN=https://oauth.xiangenhu.info, GENAI_GCS_BUCKET to a private bucket, and API_KEY_ENCRYPTION_SECRET to a stable secret of at least 32 bytes. Grant storage access to the Cloud Run service identity. Preserve the encryption secret across revisions.","Preserve all existing OAuth/SMTP environment values, SERVICE_KEYS and CORS configuration. This branch does not enforce REDIRECT_URI_ALLOWLIST; if an external policy enforces callback restrictions, permit the canonical /genai/auth/callback there. App callbacks are registered separately in GenAI Administration. ADMIN_SUBJECTS is preferred for admin identity; ADMIN_EMAILS is available for bootstrap.","HOSTED_MODELS_JSON=[] enables personal/local usage without host keys. Hosted models require separately configured host credentials, operator-approved tariffs and allowances. Do not copy example prices as current vendor prices.","Acceptance checks: existing app login and SMTP still work; linking rejects wrong/replayed state; keys never reach the browser or app backend; saved connection test succeeds; non-authorized connections are rejected; streamed requests finish with usage or explicit uncertainty; revocation works; refreshes serialize; duplicate IDs never trigger a second generation.","Initial limits: 20 saved connections, 100 grants and 2,000 operations per account/month, five pending/uncertain calls, plus per-instance AI admission. This GCS-backed release is not certified for millions of active users. Normalize the operational datastore and perform production-like load testing before mass rollout.","Teacher/class sponsorship, online credit checkout, xAPI export and native enterprise IAM are not implemented in this addition. Keep these as explicit future work rather than inferring support from the provider templates."]},{"id":"agent","title":"13. Handoff for an AI coding assistant","text":["Give the assistant the plain-text version of this page and the SDK. Ask it to integrate only the GenAI connection flow and request adapter, preserving your existing login and email behavior."],"code":"Integrate the SKO GenAI gateway at https://oauth.xiangenhu.info.\nRead https://oauth.xiangenhu.info/genaiInstruction.txt and use the server-only SDK at\nhttps://oauth.xiangenhu.info/genaiInstruction/sdk.js.\nPreserve existing OAuth login, user identities, permissions and SMTP behavior.\nAdd Connect/Manage AI to profile settings with explicit per-app user consent.\nKeep app credentials and user tokens encrypted on the backend.\nImplement one-use state/PKCE linking, serialized token refresh, scoped connection\nselection, durable request IDs, streaming and uncertain-outcome handling.\nDo not use a gateway login JWT as a GenAI access token. Do not silently fall back\nto a host key, retry inference with a new ID, or invent unsupported API features.\nUse the app's durable session/token stores and include targeted integration tests.\nReport runtime configuration still required; do not claim deployment from a commit."}],"availability":"Public documentation does not prove GenAI is enabled or deployed with working credentials."}