How to Use LLM APIs: Tokens, Costs, Keys and Your First AI Feature
What an LLM API is, how tokens, context windows and per-token pricing work, how to keep your API key safe, and how to add a first AI feature to your app.
Mythex Team · · 7 min read
An LLM API is a web service that lets your app send text to a large language model — such as OpenAI's GPT models, Anthropic's Claude or Google's Gemini — and get its reply back in code. You sign up with the provider, create an API key, and your app's server sends requests with that key; you're billed for each request by the number of tokens (chunks of text) going in and coming out. That's all "adding AI to an app" usually means: a well-written request to someone else's model, made safely from your server.
Why it matters when you build with AI
There are two different AIs in an AI-built app, and it's easy to mix them up:
- The AI that builds your app — the app builder's own agent, which writes the code.
- The AI inside your app — a feature your users use, like a chat assistant, a "summarise this" button or automatic tagging.
The second one calls an LLM API with your key and bills your provider account. Understanding tokens, context and pricing is what lets you design a feature that works, doesn't leak your key, and doesn't surprise you with a bill.
An everyday analogy: a very capable freelancer, paid by the word
Think of an LLM API as a brilliant freelance writer you reach by email and pay by the word — both for the words you send them and the words they write back (and the words they write cost more).
- They have no memory between emails. If you want them to continue a conversation, you paste the earlier messages into each new email.
- They can only read so much at once. That limit is the context window.
- The clearer your brief, the better and shorter the answer — which is also cheaper.
The core concepts
Tokens
Models read and write text in tokens. Anthropic's pricing documentation estimates about 4 characters, or roughly 0.75 English words, per token as of September 2026. Different models and languages tokenise differently, so treat that as a rough guide. Everything counts: your instructions, the user's message, any documents you include, earlier conversation, and the reply.
Input and output
Input tokens are what you send. Output tokens are what the model writes. On the pricing pages of OpenAI, Anthropic and Google, as of September 2026, output tokens cost several times more than input tokens for the same model. Google's Gemini pricing also notes that a model's internal "thinking" counts as output.
Context window
The maximum number of tokens a model can take in a single request, including the reply. Modern models accept very long inputs — some of Anthropic's current models, for instance, accept up to a million tokens. But long inputs cost more and take longer, so send only what the task needs.
Stateless requests
With most chat-style APIs, each request stands alone by default. To make a chat "remember", your app sends the earlier messages again with every new one. That's why a long conversation costs more per message than a short one.
The system prompt
Instructions that set the model's role and rules for every request: "You summarise support tickets in three bullet points. Never include customer email addresses." Users don't see it, but it shapes every reply.
Streaming
Instead of waiting for the full answer, the API can send text as it's generated, so users see words appear immediately. It makes chat interfaces feel much faster.
How pricing works
As of September 2026, checked on the providers' own pricing pages:
| Pricing lever | OpenAI | Anthropic | Google (Gemini API) |
|---|---|---|---|
| Unit | Per million tokens | Per million tokens | Per million tokens |
| Input vs output | Priced separately; output higher | Priced separately; output higher | Priced separately; output higher |
| Reused (cached) input | Discounted | Discounted cache reads; writing to the cache costs a little more | Discounted, plus a storage charge |
| Batch (non-urgent) requests | 50% off | 50% off | 50% off |
Some providers also offer a way to try before paying: Anthropic gives new users a small amount of free credit, and Google's Gemini API has a free tier on some models, with the caveat that content sent on it may be used to improve Google's products.
Prices differ a lot between a provider's small, fast models and its largest ones, and they change often, so we don't quote them here. A simple estimate for any feature is: (input tokens × input price + output tokens × output price) × number of requests per month. Then add a margin, and set a monthly spending limit in the provider's dashboard.
A worked example: a "summarise this ticket" button
A small support team wants a button that turns a long customer email thread into three bullet points and a suggested reply.
- Get a key. Create an account with one provider and generate an API key. Set a low monthly spending limit.
- Store the key as a secret on your server, named something like
ANTHROPIC_API_KEYorOPENAI_API_KEY— never in the page's code. See environment variables and secrets. - Add a server route,
POST /api/summarise, that:- checks the person is logged in and belongs to the support team;
- loads the ticket from the database;
- sends the provider a system prompt ("Summarise in three bullets, then draft a polite reply; don't invent order numbers") plus the ticket text;
- caps the reply length with a maximum output tokens setting;
- returns the result and logs how many tokens it used.
- The button calls your route, not the provider, and shows the summary while it streams in.
- Handle failure. If the provider is slow or returns an error, show "Couldn't summarise right now" and keep the ticket usable.
- Measure. After a week, look at the token logs. If most of the cost is long email threads, trim quoted history before sending.
Notice what's missing: no model training, no special infrastructure. One key, one server route, one well-written prompt.
Key terms
- API key: the secret that identifies your account to the provider. Anyone who has it can spend your money.
- Model: the specific LLM you call. Providers offer several sizes at different prices and speeds.
- Temperature: a setting for how varied the replies are. Lower is more predictable.
- Max tokens: a cap on reply length, which also caps cost.
- Rate limit: how many requests or tokens your account may use per minute.
- Embeddings: numbers that represent the meaning of text, used for search over your own documents.
- RAG (retrieval-augmented generation): finding relevant passages from your own data and including them in the prompt, so the model answers from your content.
- Tool use / function calling: letting the model ask your app to run a specific function, like "look up order status".
Common mistakes and misconceptions
- Putting the key in frontend code. The most expensive mistake. Browser code is public. Always call from your server. The security checklist for AI-built apps covers this with other launch checks.
- No spending limit or per-user limits. A bug, a loop or an abusive user can run up costs quickly. Set a provider limit and limit requests per user in your app.
- Sending everything every time. Whole documents and entire chat histories multiply input tokens. Trim, summarise older messages, or retrieve only relevant parts.
- Trusting the output blindly. Models can state wrong things confidently. For facts, prices or anything legal or medical, ground answers in your own data and show sources, or have a person check.
- Letting user input override your rules. Users can type "ignore your instructions". Keep permissions and sensitive actions in your server code, not only in the prompt.
- Starting with the biggest model. Try a smaller, cheaper model first; move up only if quality demands it.
- Sending personal data without thinking. Check the provider's data-use terms, especially on free tiers, and send only what the feature needs.
What to ask your AI builder for
Add a "Summarise" button on the ticket page. It calls a server route that uses the secret
OPENAI_API_KEY(never exposed to the browser), checks the user is logged in, sends the ticket text with this system prompt: […], limits the reply to 300 tokens, streams the answer to the page, and shows a friendly error if the call fails. Log tokens used per request, and limit each user to 30 summaries an hour.
For a full chat experience, see how to build an AI chatbot app; for other ideas, how to add AI features to your app. If APIs in general are new to you, start with what an API is.
LLM APIs in apps built with Mythex
When you build with Mythex, the AI inside your app uses your own provider key. You add it as a project secret, or through the secrets card Mythex shows in chat when it needs one, and ask the agent to make the call on the server. Provider usage is billed to your provider account, not to Mythex build credits. Mythex doesn't include tokens for your app's users or a built-in knowledge-base product, and rate limits and cost controls are yours to design and test. The docs recipe add AI features with your own API keys has a starter prompt, and the AI app builder page shows where to begin.
Questions
What is an LLM API?
An LLM API is a web service from an AI provider such as OpenAI, Anthropic or Google that lets your app send text (and often images or files) to a large language model and get the model's reply back. You pay per use, measured in tokens, and authenticate with an API key.
How much does an LLM API cost?
As of September 2026, OpenAI, Anthropic and Google all price their main models per million tokens, with separate rates for input and output, and output costing more. The total depends on the model, how much text you send and how long the replies are. Check each provider's pricing page, and set a spending limit.
What is a token?
A token is a chunk of text the model reads or writes, often part of a word. Anthropic's pricing documentation estimates about 4 characters, or roughly three-quarters of an English word, per token; exact counts vary by model and language.
Can I call an LLM API directly from the browser?
You shouldn't with a secret key. Anything in browser code can be read by visitors, who could then use your key and run up your bill. Send requests from the browser to your own server, and have the server call the provider.
Do I need to train a model to add AI to my app?
Almost never. Most AI features use an existing model through an API, guided by clear instructions in the prompt and, where needed, your own data included in the request. Training or fine-tuning is a later optimisation, not a starting point.