AI tools
Call OpenAI, Anthropic or Google from your app through a named AI tool, priced in credits.
A tool's definition lives on the server: its instructions, prompt template, inputs, model, caps, credit price and the access it requires. Your provider key stays on the server and never reaches the browser. A signed-in person's app sends the tool's name and inputs. Gemmein composes the request, spends the tool's credits, adds your key and streams the answer back.
Call a tool from your app
A tool file at gemmein/ai/tools/deep-research.json:
{
"label": "Deep Research",
"provider": "openai",
"model": "gpt-4o",
"credits": 20,
"requires": "access:pro-max",
"instructions": "You are a careful research assistant. Answer with sources.",
"promptTemplate": "Research this for a {{audience}} reader:\n\n{{question}}",
"inputs": [
{ "name": "question", "type": "text", "required": true, "maxLength": 2000 },
{ "name": "audience", "type": "text" }
],
"bounds": { "maxOutputTokens": 4000 }
}
The app never sends a prompt or a provider request. It names the tool and passes the inputs the tool declares. Gemmein composes the request in the provider's own format, from the instructions and the template on the server. The answer comes back in the provider's own shape, so a streaming call streams:
const res = await g.ai.run("deep-research", { question: text, audience: "beginner" }, { stream: true })
for await (const chunk of res.body) render(chunk) // the provider's SSE, byte for byte
res.headers.get("x-gemmein-credits-remaining") // the balance after this call
What it is. A named AI operation, defined on the server, that you price in credits and restrict by access.
Does.
- Composes the provider request from the tool's own instructions and template, with the inputs the app sent.
- Runs it on your key with the pinned model, or the model alias the tool names, and the output cap.
- Charges the tool's price, per call or per tokens (see Pricing).
- Records every call on
ai_calls(who, tool, tokens, units and unitCount, credits, outcome). It keeps the prompt and answer only when you switch that on for the tool.
Does not. Does not let the browser compose the request or see the prompt (raw calls are off unless the owner switches them on for a provider key); does not let the browser set a price, a model or the access it requires; does not run from a secret key; does not rename a tool after creation.
Needs something else when. You want a plan-dependent price for the same
operation → make two tools, each with its own access requirement; you meter something that is not an AI call
→ spendCredits from your server; you need an image, audio, video or a
transcript → a job tool and a run; you need embeddings → your server
calls the provider directly.
Example. "Deep Research", openai, gpt-4o, 20 credits, requires access:pro-max, with
instructions and a question input.
For a non-stream answer as one string, whichever provider answered, and for the person's own history:
const answer = await g.ai.runText("deep-research", { question: text })
const { calls } = await g.ai.calls() // what they ran, when, what it cost, how it ended
g.ai.runText reads choices[0].message.content from OpenAI, joins
content[].text from Anthropic, and joins
candidates[0].content.parts[].text from Google.
A provider's non-2xx answer is returned as it came by g.ai.run, and is not
thrown. g.ai.runText, like g.ai.text, throws it as
provider_error with the provider's status and message. A refusal by Gemmein
throws GemmeinError with one of the codes below. An unknown
input, a missing required one, a wrong type or a value over its cap is refused by name before
anything is spent (invalid_inputs).
What the person sees:
| When | Message |
|---|---|
| They lack the access | "Deep Research requires Pro Max." |
| They are short of the price | "Deep Research costs 20 credits. You have 7." |
| They are short of a per-token tool's ceiling | "Deep Research needs up to 12 credits. You have 7." |
| The ledger line after the call | "20 credits spent · Deep Research · 87 remaining." |
Tool kinds and providers
A tool is one of two kinds. A chat tool answers now. A job tool (generate or transcribe) answers with a run. A provider serves every kind its key sells:
| Kind | Answers with | Providers |
|---|---|---|
| chat (the default) | An answer now | openai, anthropic, google |
| generate | A run: an image, a video or audio out | openai, google, replicate, fal, elevenlabs, runway, external |
| transcribe | A run: audio in, text out | openai, google, deepgram, replicate, fal, assemblyai, elevenlabs, external |
An OpenAI key makes chat, image, speech, transcription and video tools. A Google key makes
chat, image, video and audio-transcription tools. external is
your own pipeline.
When your own server makes the model call, or the call is neither chat nor a job (such as embeddings), call the provider directly with your key. Ask Gemmein's server methods in the SDK reference who the person is and what they hold. If your provider is not on the list, write to hello@gemmein.com.
Files a tool takes
A tool says which kinds of file a person may hand it: image, audio,
video and document. Each is a tick in the tool's file
("accepts": ["image", "document"]) and on the dashboard's AI page. A
file input takes the ref of a file the person already uploaded, to any
collection: g.ai.run("describe-photo", { photo: upload.ref }).
Gemmein checks the file's kind, the tool's model and the file's size before any credit moves, then hands the file to the provider. Most providers fetch a short-lived signed link made for that one call. Small files ride inline. Gemini files over 14 MB stream from storage into Gemini's own file store. A file is never sent as a public link, never converted, and never held whole in Gemmein's memory.
A tick is offered only where the tool's provider and model read that kind:
| Provider | Takes | Models |
|---|---|---|
| image, audio, video, PDF | Gemini 1.5, 2.x and 3.x | |
| openai | image · audio | image: gpt-4o, gpt-4.1, gpt-5, o1, o3, o4-mini · audio: gpt-4o-audio models (mp3 or wav); transcription: whisper-1, gpt-4o-transcribe |
| anthropic | image · PDF | image: Claude 3 and later · PDF: Claude 3.5 Sonnet and later |
| assemblyai, deepgram, elevenlabs | audio, video | transcription models (ElevenLabs Scribe) |
| replicate, fal | image, audio, video | models that take a file link |
| runway | image | image-to-video models |
| external | all four | your own pipeline, by a short-lived signed link |
npx gemmein dev and npx gemmein sync refuse a tick the model
does not read and name the models that do. A file may be up to 500 MB of video, 100 MB of
audio, or 25 MB of image or document, or the provider's own limit when lower.
"inputMaxMb": { "video": 200 } in the tool's file sets a lower limit for one tool;
the AI page has no size field. A tool that takes files is priced per run: a flat credit price
per call, held and charged as that price whatever the file's size. A video input can cost
more on your own provider key. A tool charged per token that ticks a kind is refused at
gemmein dev, gemmein sync and on the AI page: "A tool that takes files
is priced per run — set a credit price per call, and lower its file size limit if you want to
cap your cost." Text-only tools keep per-token pricing. Files a person hands a tool count
toward your app's storage; video fills it fastest, and you can add 100 GB blocks under
Usage. With ticks set, a file input takes only a Gemmein file ref, never a link. A tool
created before ticks takes exactly the files it took before.
Create a tool
Create a tool on the dashboard's AI page, or have your AI write a file at
gemmein/ai/tools/<name>.json. npx gemmein sync pushes it to
Development. npx gemmein sync --live pushes it to your live app with a sync
key. Create the sync key under Live → Secret keys → Sync key with your sign-in code.
It works for one hour, is shown once and is never saved.
The file owns the implementation on every sync. After creation, the dashboard owns the commerce: label, credits, unit and per, access, on/off, record calls, and a transcribe tool's delete-after-run. Every field is in the field table.
A tool can be defined before its provider key exists. Runs are refused
ai_not_configured until the key is added on the AI page, and the dashboard shows
the tool as waiting on a key until then. A tool's name is fixed once created, because it is
what the app calls.
Model aliases
What it is. A name of your own, such as fast or writer,
that stands for one provider and one model. A tool names the alias in place of a model.
Does.
- Resolves at call time to the alias's provider and model. Repointing the alias moves every tool on it, with no edit to any tool and no deploy.
- Lives on the AI page or in
gemmein/ai/aliases.json(one object of name → { provider, model }), and syncs with the tools. - Refuses at save a model the provider key's allowlist does not admit.
- Keeps a tool and its alias on one provider.
Does not.
- Span providers. An alias is on one provider, and a tool on another cannot name it.
- Exist on
external(your pipeline has no model). - Change a tool's price, access or inputs.
- Sit beside a pinned model. A tool names a model or an alias, never both.
- Stop an alias being removed while tools still name it. Such a tool answers
409 tool_incompleteuntil it is repointed.
Needs something else when.
- You want a different model per plan → two tools, each with its own access requirement.
- You want the app to choose the model → not offered; the model is the server's.
- You want a model the allowlist lacks → add it to the provider key's models on the AI page first.
Example. fast → openai, gpt-4o-mini. "Quick answer" and
"Summarise" both name fast. Move fast to gpt-4.1-mini
on the AI page and both tools run on it from their next call.
Provider keys
The AI page holds one key per provider, per environment. Paste it, make a test call, and it is set. A key is 8 to 512 printable characters. Replace it by pasting again. A job provider's key is pasted like a chat key; its test call is one authenticated read that spends nothing.
The key is write-only. No response, no audit record and no page shows more than its last
four characters. The AI page shows the last four characters of a key of 16 or more characters
(a shorter key shows dots) and the last call. A provider that echoes the key in a refusal
reaches the app as ***<hint>.
Removing a key is a step-up action, and your app's AI features stop the moment it is removed.
Optionally, list up to 20 models the app may call; the test call uses the first. Any other
model answers model_not_allowed. The model name is the provider's, 1 to 80
characters of letters, digits, dots, colons, underscores and hyphens. Google needs one;
OpenAI and Anthropic refuse a missing one themselves.
Pricing
A tool holds one credit price. For a price that depends on the plan, make two tools, each with its own access requirement. How credits are bought, granted and spent is on Credits.
Per call
call is the default unit. The tool's credits are spent before the request is
forwarded; the default tool spends one. A person short of the price is answered
402 credits_exhausted before anything is sent, with the balance in the
message.
Per 1,000 tokens
With unit: "tokens", per is 1,000 by default. Before the call the
tool reserves a ceiling: the price of the request's bytes ÷ 3 plus the output cap (4,096 when
the tool sets none). When the answer ends cleanly, the call settles at the provider's usage,
up to the ceiling, and what was not used returns. Per-token pricing is for text: a tool that
takes files is priced per call.
The call costs the ceiling when the provider reports no usage, when a stream does not end cleanly, or when the call is abandoned mid-answer. A call abandoned before the provider answered costs the input estimate.
Per job unit
A generate tool meters images, seconds or characters
(or call: one price per run). A transcribe tool meters seconds (or
call). A tool's unit is the one its model meters: an image model per image, a
video or transcription model per second, a speech model per character, or call.
Any other unit is refused when the tool is saved, since a unit nothing meters would charge every person the
ceiling.
A metered job tool reserves credits × ceil(bounds.maxUnits / per) when its run
starts, and settles at what the provider metered when the run succeeds.
Refunds
Credits spent before the request are refunded only when the provider fails before its
first byte, with a non-2xx answer or no answer at all. The response then carries
x-gemmein-credit: refunded. A stream that dies after the first byte is not
refunded, and hanging up early does not refund.
Response headers
Every answer that passed the spend carries x-gemmein-credits-remaining.
x-gemmein-tool names the tool; absent on the implicit default.
Usage & billing
The Usage & billing page counts AI calls for the last 30 days: spent from your people's credits, priced by your provider. Gemmein meters the calls; your provider bills the tokens.
Image, audio and video runs
A job tool (kind generate or transcribe) answers with a
run. The app calls the same endpoint, POST /ai/run/<tool>, and gets
202 with the run straight away. Gemmein follows the run until it ends.
What it is. One job on a provider that runs jobs (see Tool kinds and providers) or on your own pipeline, started by a signed-in person on a job tool. The server holds it from start to end: its status, its progress, its price and its outputs.
Does.
- Reserves the tool's ceiling when the run starts:
credits × ceil(maxUnits / per), or the price itself on a per-run tool. A refusal names it ("Poster needs up to 8 credits. You have 4."). - Composes the provider request from
paramsand the inputs, on your key, and submits it. - Checks every open run every 10 seconds, polling the provider at 5 seconds, then 15, 30 and 60.
- When the run succeeds, copies every output URL the provider answered into a sealed
file on the person: image, audio or video, up to 200 MB each, with the type proven from
the bytes. A transcript comes back as
text. - Settles the hold at what the provider metered (the count of images, the provider's own seconds rounded up, or the request's characters) at the tool's rate, and never above the hold.
- Ends an open run
cancelledong.runs.cancel, andexpired60 minutes after it started, and tells the provider to stop where it can be told. - Records every ended run on
ai_calls, and names the run on the ledger's reserve and settle entries.
Does not.
- Apply a webhook body. A provider's webhook only makes Gemmein check the run sooner, and Gemmein fetches the run again by the provider's own id.
- Charge a failed, cancelled or expired run. The hold returns whole.
- Hand a provider a file ref. A ref becomes a short-lived signed link.
- Let a run bill past
maxUnits.
Needs something else when.
- You want the answer in the same request → a chat tool.
- Your own server calls a provider → call the provider directly, with your key.
Example. "Poster", replicate, 2 credits per image, up to 4 images a run, with a
prompt input.
The tool file, at gemmein/ai/tools/poster.json. The count
placeholder stands alone in its string, so the provider receives a number:
{
"label": "Poster",
"kind": "generate",
"provider": "replicate",
"model": "black-forest-labs/flux-schnell",
"credits": 2,
"unit": "images",
"params": { "prompt": "{{prompt}}", "num_outputs": "{{count}}", "aspect_ratio": "16:9" },
"inputs": [
{ "name": "prompt", "type": "text", "required": true, "maxLength": 1000 },
{ "name": "count", "type": "number" }
],
"bounds": { "maxUnits": 4 }
}
Starting it holds 8 credits. Two images back settle at 4 and return 4. In the app:
const run = await g.runs.start("poster", { prompt: "a lighthouse at dusk", count: 2 });
const done = await g.runs.watch(run.id, { onUpdate: (r) => setProgress(r.progress) });
if (done.status === "succeeded") img.src = done.result.files[0].url; // a signed link, created for this answer
save(done.result.files[0].ref); // the ref is what the app stores
A result file's url lapses at urlExpiresAt. Read the run again for
a fresh one. The ref stays, and is what g.files reads. The run carries
reserved, and once ended charged, units and
unitCount.
A key on start (your own name for the run, up to 80 characters)
makes a retried start of the same tool answer the same run. g.runs.list({ since })
is the person's own runs, newest first; g.runs.get(id) is one of them.
Runs are the person's, called from the browser. A secret key is refused, and another
person's run is unknown_run. Ten runs may be open per person at a time.
| Status | Meaning |
|---|---|
queued | Created, the ceiling reserved, the provider not yet asked |
running | The provider has it; progress is 0–100 when the provider says |
succeeded | Ended well; result is set and the hold settled at what was metered |
failed | The provider gave up; error says why; the hold returned |
cancelled | g.runs.cancel ended it; the hold returned |
expired | 60 minutes passed with no end; the hold returned; the provider told to stop |
Run refusals are in the codes table. start also meets
every refusal g.ai.run does for the tool itself: unknown_tool,
tool_disabled, entitlement_required,
payload_too_large, ai_capped, session_required,
scope_denied.
Your own pipeline
What it is. A job tool whose provider is external. The work
runs on your compute, at a URL you name, in place of a provider Gemmein calls for you.
Does.
- Starting a run POSTs it to your URL, signed, on Gemmein's own retry schedule, the same way a job goes to Replicate.
- Needs no provider key: the provider-key endpoint refuses
external, because your pipeline is the provider. No alias, andmodelis your own label or nothing. - Keeps the run on Gemmein: its reservation, its person, its deadline, its result. No relay row is created for it.
- Works with anything behind that URL: a LangGraph graph, a CrewAI crew, a multi-agent system, a person in the middle. It counts as one tool, with one run and one price.
- Sends it the run signed with the person, the inputs, a credit ceiling and a deadline of up to 24 hours. The work runs on your own servers.
Does not.
- Need a relay or create a relay row.
run_startedfires react-only if you have one, but nothing requires it. - Hand your pipeline a provider key or a model. There is none.
- Bill past
bounds.deadlineMinutes(1 to 1,440; 60 when unset; only an external tool may set it). Past it the run expires and the hold returns, whatever your pipeline is doing. - Accept a second
complete, acompleteafterfail, or a call past the deadline. The token completes once.
Needs something else when.
- Your work finishes inside the request → a chat tool, or a job on a provider Gemmein already calls (Replicate, fal, …). No pipeline needed.
- You want the tool free for subscribers →
credits: 0withrequiresAnyset (a priced-or-included tool; seecreditsin the field table).0with norequiresAnyis refused at save.
Example. A weekly report generator: url points at your own function,
bounds.deadlineMinutes: 30, credits: 0 with
requiresAny: ["access:pro", "access:max"]. It is included with either plan,
still metered on the ledger, never charged.
A research crew (a planner, three researchers and a writer agent, wired with CrewAI) is
the same tool. url points at the crew's own endpoint,
bounds.deadlineMinutes: 120, and it completes with the written report in
text once every agent in the crew has run (run files hold images,
audio and video).
A start with no url set answers 409 tool_incomplete ("… names no
URL — set where it runs") before any credit moves.
Agents
A relay's start_run starts this tool, or any tool, for that person when a
record is written, a payment lands, a clock fires or a webhook arrives, with nobody pressing a
button. run_started chains it to the next tool, one hop at a time, in a sequence
your app composes. Gemmein owns no graph and runs none of the agent's code. It holds the run's
reservation, person, deadline and result, and nothing else.
The hand-off
Gemmein signs the hand-off the same way it signs a relay's call_url. The two
envelopes are what Gemmein sends your URL and what your pipeline sends back:
// Gemmein → your URL (X-Gemmein-Signature: t=…,v1=…)
{
"run": { "id": "run_…", "tool": "poster", "kind": "generate", "units": "images", "reserved": 8, "createdAt": "…" },
"params": { "prompt": "…" }, "person": { "id": "usr_…", "email": "…" }, "deadline": "…",
"complete": { "url": "…/complete", "token": "…", "expiresAt": "…" },
"progress": { "url": "…/progress" }, "fail": { "url": "…/fail" }
}
// your URL → Gemmein (Authorization: Bearer <token>)
{ "text": "made: a lighthouse at dusk", "units": 2 }
Answer a run
Your pipeline answers with two required calls and one optional one, all with the token the hand-off carried:
| Endpoint | Body | Answers |
|---|---|---|
POST …/complete | { units?, files?: [{ url, contentType? }], text?, data? }, up to 64 KB | 200 { run }: succeeded, settled at units, never above the hold |
POST …/fail | { error? }, up to 2,000 characters | 200 { run }: failed, the hold returned |
POST …/progress (optional) | { progress: 0..100, note? }, any number of times | 200 { run }: non-spending, never takes the lease |
The pipeline helper
g.pipelines.handle does all of this in three lines (@gemmein/sdk
≥ 0.12.0):
export default g.pipelines.handle(process.env.PIPELINE_SECRET!, async (run, { progress }) => {
await progress(50, "half way");
return { text: `made: ${run.params.prompt}` };
});
g.pipelines.handle(secret, fn, options?) verifies the signature, acknowledges
within Gemmein's 10-second window, runs fn in the background, and calls
/complete or /fail for you. It is a plain Fetch-API handler that runs
unmodified on Vercel, Cloudflare Workers, Deno, Bun or Node 18+.
secret is the tool's own signing secret (aisk_…), created with the
tool, shown once on the tool's page and rotatable. It is never your app key. A bad or stale
signature is refused before fn runs. A re-delivered hand-off (the same
run.id, after a lost acknowledgement) runs fn once.
Local development
gemmein dev answers AI tools without a key: a fake provider echoes a stream,
marked x-gemmein-ai: fake, and the tool's credits are spent from the local
ledger. For a real call, set GEMMEIN_AI_KEY_OPENAI,
GEMMEIN_AI_KEY_ANTHROPIC or GEMMEIN_AI_KEY_GOOGLE in the environment
gemmein dev runs in. Keys never sync; your hosted app holds its own, pasted on the
AI page.
Tools do sync. A file at gemmein/ai/tools/<name>.json hot-reloads like a
relay, and the boot card prints a line per tool: AI tools: deep-research (20 credits ·
requires access:pro-max · openai).
Credits in Development
A fresh person's balance is zero, and AI tools refuse at zero locally and on Gemmein alike,
so testing starts with a credit arriving. Locally that is the pay simulator. Declare a
product with "Grants credits" in gemmein/payments.json
(npx gemmein payments setup asks for it), call
g.payments.buy("credits pack"), and open the URL it returns. The
simulated payment travels the real signed-webhook path, and the credits land on that
person's local ledger.
A relay's grant_credits is the other local way, for a receiver you have wired.
On Gemmein the two ways are the same purchase, plus a comp you add by hand on that
person's page under Customers. There is no comp in
gemmein dev, because the dashboard is not part of it.
Runs and pipelines in Development
On gemmein dev without a key for the provider, a run ends succeeded with a
generated image or a placeholder transcript. It is metered and settled, so the whole flow
works before any key exists.
An external tool works the same way on gemmein dev, with loopback allowed.
Your pipeline calls
http://127.0.0.1:<port>/runs/<app>/<env>/<id>/complete with the token.
What happens if…
| If… | What happens | What you do |
|---|---|---|
| The provider fails before its first byte, with an error or no answer | The person gets the tool's full price back. The response carries x-gemmein-credit: refunded and the provider's own status | Nothing; the person can retry |
| A stream stops mid-answer, or the person hangs up | No refund: the request is already on your provider bill | Nothing |
| The person has fewer credits than the tool costs | 402 credits_exhausted, naming the tool, its price and their balance. The provider is never called | Show the credit pack |
| One person makes more than 20 calls in a minute | The next call answers 429 ai_capped with resetAt and spends nothing. The limit is per person; other people are unaffected | Wait for resetAt |
| You create a tool before pasting its provider key | The tool saves, and the AI page shows it waiting on a key. Calls answer ai_not_configured until the key is there | Paste the key on the AI page |
| A tool requires a plan the person lacks | 403 entitlement_required, naming the plan. If that plan was deleted, the person reads a neutral sentence and your Logs row keeps the detail | Sell or grant the plan |
| You delete a tool | Calls naming it answer 404 unknown_tool and spend nothing. Runs already open still end, failed, with their holds returned | Nothing |
| A call is refused in Live | The person reads a neutral sentence. Your Logs row names who, which tool and why, and Problems groups it | Read the Logs row |
| A provider meters more than the tool's ceiling | The run is charged its hold, never more | Raise bounds.maxUnits if the ceiling is too low |
| The provider is busy (429) or down (5xx) while a run is in progress | Gemmein checks again later, until the run's deadline. A job is never submitted twice | Nothing |
| The provider refuses or fails the submit itself | The run fails, says why and returns its hold. It is not sent again | Start a new run |
A run is still open 60 minutes after it started (or its own deadlineMinutes) | It ends expired, its hold returned, and the provider is told to stop | Nothing |
| The person cancels a run | It ends cancelled, its hold returned, and the provider is told. A second cancel answers 409 run_ended | Nothing |
| Your pipeline answers the hand-off with a 4xx other than 408 or 429 | The run fails and its hold is returned. Any other failure is retried, with the same run.id, until the deadline | Fix the URL or the pipeline |
| A person hands a tool a kind of file it does not tick | 400 input_not_accepted, naming the kinds it takes. Nothing is spent | Send a kind it takes, or tick the kind on the AI page |
| The tool's model does not read that kind (an alias was repointed) | 409 input_not_supported. Nothing is spent | Pick a model that reads it, or untick the kind |
| The file is bigger than the tool's, the provider's or the platform's limit | 413 input_too_large, naming the limit. Nothing is spent | Send a smaller file |
| The file is sealed and belongs to someone else, or was deleted | 400 invalid_inputs: file not found. Nothing is spent | Send a ref of a file this person uploaded |
| A person sends a second large file while their first is still streaming to the provider | A chat tool answers 429 input_busy and spends nothing. A run waits and tries again | Retry when the first finishes |
| A tool that takes files costs more than the person holds | 402 credits_exhausted, before any file is read | Show the credit pack |
| A tool charged per token ticks a file kind | Refused at gemmein dev, gemmein sync and on the AI page: a tool that takes files is priced per run | Set a credit price per call |
| Your app's storage is full (past twice its ceiling) | Uploads are refused with 429 usage_limit_exceeded and a clear message. Files already stored stay | Delete files, or add a 100 GB block under Usage |
| The provider refuses the file | A chat tool answers 502 input_rejected and spends nothing. A run fails with the provider's sentence and returns its hold | Check the file, then retry |
A hand-off reaches g.pipelines.handle with the wrong secret, an altered body or an old timestamp | Refused before your function runs. The same run.id delivered twice runs your function once; a function that throws fails the run and returns its hold | Check the pipeline secret |
Reference
Tool fields
| Field | Rule |
|---|---|
name | 1–40 lowercase letters, numbers or hyphens, starting with a letter or number. Fixed once created |
label | 1–60 characters, shown to the person and on the ledger |
kind | chat (the default: answers now) or a job kind: generate (an image, a video or audio out) or transcribe (audio in, text out). A job tool answers with a run |
provider | one that serves the tool's kind (see Tool kinds and providers); a tool on a provider that does not serve its kind is refused at save. A job tool may instead run on external, your own pipeline: no key, no alias, and model your own label (up to 120 characters) or nothing |
model | pins the provider's own model id (gpt-4o; owner/name on Replicate, fal-ai/flux/dev, gen4.5). On openai and google a job tool's model names the endpoint: gpt-image-1 (images), gpt-4o-mini-tts (speech), whisper-1 or gpt-4o-transcribe (transcription), sora-2 (video); imagen-4.0-generate-001, veo-3.0-generate-001, gemini-2.5-flash (audio in). Deepgram's is nova-3; AssemblyAI takes none |
accepts | the kinds of file the tool takes: any of image, audio, video, document, each one the provider and model read (see Files a tool takes). Absent: as before ticks (a chat tool takes no files; a job tool takes what its provider takes) |
inputMaxMb | a lower limit per ticked kind, in whole MB: { "video": 200 }. Never above the provider's or the platform's limit. Optional, set in the tool's file only |
url | provider external only: your pipeline's own address (https; loopback http:// only on gemmein dev). Set on creation or later; a run started before it is set is refused tool_incomplete before any credit moves. See Your own pipeline |
credits | a whole number from 1 to 10,000, or 0 only with requiresAny set: the tool is included with access, still metered, never charged. 0 with no requiresAny is refused at save |
unit / per | call (the default; per is 1) or tokens (per 1,000 by default, 1 to 1,000,000). A generate tool meters images, seconds or characters, a transcribe tool seconds, and either may use call. How each unit reserves and settles is under Pricing |
params | a job tool's provider request, composed on the server in place of promptTemplate: a JSON object whose string values may carry {{input}} placeholders naming declared inputs ({ "prompt": "{{prompt}}", "aspect_ratio": "16:9" }). A string that is one placeholder and nothing else takes the input's typed value, so a number stays a number. The browser never shapes the request |
alias | the name of one of the app's model aliases, in place of model, never both |
requires / requiresAny | requires is one access key (access:pro-max); requiresAny is up to 10 plan or product keys. The person needs any one of them, so a Max plan may use a tool that requires Pro without a second grant. requiresAny wins when both are read; a tool that still names only requires behaves as requiresAny with that one key. A person holding none is refused, naming every key |
instructions | the system prompt, up to 20,000 characters. Never leaves the server |
promptTemplate | the user turn, with {{input}} placeholders naming declared inputs; absent, the inputs are rendered one per line |
inputs | up to 20 of { name, type: text | number | boolean | file, required?, maxLength? }; a text input holds 4,000 characters unless it says (up to 20,000). A file input takes a sealed file ref the person already owns (what g.files.upload returned for an image or document, or an earlier AI run's own output for audio or video) or an https URL. The provider receives a short-lived signed link, never the ref |
recordCalls | keep the prompt and answer on ai_calls. Off by default |
bounds | maxOutputTokens (a composed call's ceiling; 4,096 when unset), stream, maxBodyBytes (raw calls only, up to 262,144), deadlineMinutes (external tools only: how long a run may take before it expires, 1 to 1,440; 60 when unset) |
bounds.maxUnits | a metered job tool's ceiling: the most units one run may bill, 1 to 1,000,000. Required when unit is images, seconds or characters; the hold on every run is priced from it |
Numbers
- The credits the owner set for that tool; one by default.
- Up to 50 tools per environment; a name is at most 40 characters, a label at most 60.
- 20 calls per person per minute.
- 64 KB of inputs on a
g.ai.runcall. - 256 KB raw request body on a
g.ai.chatcall, unless a tool sets a smaller cap (up to 262,144 bytes); nested at most 32 levels. - 170 seconds in all; on a stream, 10 seconds to the first response headers.
- Every call counts toward the app's request band like any other request.
- One test call per minute from the AI page; a test call spends no credits.
Codes
| Code | Status | Meaning · what to do |
|---|---|---|
raw_calls_off | 403 | The browser may not compose provider requests for this provider. Call a named tool with g.ai.run, or switch raw calls on for the key on the AI page |
invalid_inputs | 400 | An input is unknown, missing, the wrong type or over its cap, or a file input is not a sealed file ref or an https URL. The message names the input |
input_not_accepted | 400 | The file is a kind the tool does not tick, or a type its provider does not read. The message names what it takes |
input_too_large | 413 | The file is over the tool's, the provider's or the platform's limit. The message names the limit |
input_not_supported | 409 | The tool's model does not read that kind of file. Pick a model that does on the AI page |
input_rejected | 502 | The provider refused the file before answering. Nothing was spent |
input_busy | 429 | This person already has a large file on its way to the provider. Retry when it finishes |
tool_incomplete | 409 | The tool composes nothing: no template and no inputs. On a job tool: no params, or its alias is gone; on an external tool: no URL |
unknown_tool | 404 | No AI tool by that name in this environment |
tool_disabled | 403 | The owner switched this tool off |
entitlement_required | 403 | The tool requires a plan or product this person lacks; the message names it |
model_pinned | 403 | This tool's model is fixed; leave model out of the body |
too_many_tools | 409 | An environment holds at most 50 AI tools |
invalid_tool | 400 | Creating or updating a tool with a bad field; the message names it |
credits_exhausted | 402 | The person's balance is below the tool's price, or its ceiling on a per-token or job tool; the message names both. Show the pack |
ai_not_configured | 409 | No key for this tool's provider on this app and environment. The owner pastes one on the AI page, and the tool then runs unchanged. The person reads "This feature isn't available right now."; Development adds the owner's sentence in ownerDetail, Live never does |
provider_required | 400 | More than one provider key is set; pass provider |
model_not_allowed | 403 | The owner's allowlist names the models this app may call; the message lists them |
ai_capped | 429 | 20 calls per person per minute; wait for resetAt |
payload_too_large | 413 | On g.ai.run: the inputs are over 64 KB; send less. On g.ai.chat: the raw body is over 256 KB; shorten the conversation you send |
invalid_body | 400 | The body is not the provider's JSON request object, or is nested deeper than 32 levels |
session_required | 401 | No signed-in person; sign in first |
scope_denied | 403 | A secret key called an AI endpoint. AI tools are for the browser; a server calls the provider directly |
provider_unreachable | 502 | The provider did not answer before the first byte. Nothing was charged; the credit is refunded. Retry |
provider_error | the provider's | Thrown by g.ai.text and g.ai.runText: the provider's own non-2xx, its message in err.message. g.ai.chat and g.ai.run return it as it came |
ai_test_capped | 429 | The AI page's test call; one a minute per app |
runs_capped | 429 | Ten open runs per person; wait for one to end, or cancel one |
unknown_run | 404 | No run by this id on this person |
run_ended | 409 | cancel on a run that already ended; read its status |
not_external | 409 | A Gemmein-executed run completes on its own, never by a call to /complete |
Raw provider calls
g.ai.chat(body), where the browser sends the provider's own request, is off
until you switch raw calls on for that key on the AI page. Without the switch it answers
raw_calls_off. With the switch on, the body is forwarded as sent and the answer
passed back byte for byte. A call that names no tool runs as the default tool for one
credit.
gemmein dev's fake provider, with no key set, keeps raw calls open; there is no
switch locally. ?provider= and ?stream=1 on the URL do what the body
fields do.
| Provider | The body you send | Where it goes |
|---|---|---|
openai | Chat completions: model, messages, stream? | OpenAI's chat completions endpoint |
anthropic | Messages: model, max_tokens, messages, stream? | Anthropic's messages endpoint |
google | generateContent: model, contents; stream: true selects the streaming form | Google's generateContent endpoint for that model |
Those three endpoints are the whole list. There is no field for a URL, and no other host is reachable through a chat call.
The ai_calls record
Every call lands as one record in the ai_calls collection, per person: the
tool, the kind, the provider and model, tokens in and out when the provider said, the credits
it cost, how it ended, the refusal code, the latency and the time. The prompt and the answer
are kept only for a tool whose record calls switch is on.
A person reads their own with g.ai.calls(). You read them by app on the Records
page and on the person's record. Erasure removes a person's records.