Guides / Prompting and shipping
How to Integrate the OpenAI API into Your App
Add the OpenAI API to your app safely: API keys and projects, the Responses API, server-side calls, streaming, spend limits, rate limits and error handling.
Mythex Team · · 5 min read
To integrate the OpenAI API, create an API key in an OpenAI project, store it as a server-side secret named OPENAI_API_KEY, and call the API from a route on your own server using OpenAI's official SDK, never from browser code. As of September 2026, OpenAI's quickstart presents the Responses API as the starting point for new projects. Set a spend limit before launch, cap reply length, limit how often each user can call the feature, and handle rate-limit errors with retries.
This guide covers setup, the request flow, prompts for an AI builder, and the mistakes that leak keys or burn money. OpenAI details are from its official documentation as of September 2026. For the general concepts (tokens, context windows, pricing), read how to use LLM APIs first.
What you can build with it
- A chat assistant inside your app (how to build an AI chatbot app)
- "Summarise this", "rewrite this" or "draft a reply" buttons
- Extracting structured data from free text: turning an email into an order, tagging support tickets
- Search and question-answering over your own content
- Image generation, speech and other media features, depending on the models available
More ideas are in how to add AI features to your app.
Setting up your OpenAI account
- Create an account and a project in the OpenAI platform. OpenAI's production best practices recommend separate projects for staging and production, so each can have its own keys and limits.
- Create a project API key. OpenAI recommends setting an expiry date and rotating keys regularly.
- Set spend alerts and a hard spend limit on the limits page. OpenAI's guide says alerts notify you when usage passes an amount, and a hard limit enforces a monthly cap.
- Add billing. OpenAI's rate-limit docs describe usage tiers: your limits rise automatically as you spend more over time.
The request flow
Browser -> your server route -> OpenAI API
<- (streamed) reply <-
- The user clicks a button or sends a message.
- The browser calls your route, for example
POST /api/summarise. - Your server checks the user is logged in and allowed to use the feature.
- The server builds the request: instructions, the user's input, and any data from your database the task needs.
- It calls OpenAI with the key from the environment. OpenAI's SDKs read
OPENAI_API_KEYautomatically. - It streams or returns the reply, and logs token usage.
OpenAI's official SDK packages are called openai for both JavaScript/TypeScript (npm) and Python (pip).
Choosing a model
OpenAI offers several models at different prices and speeds, and the line-up changes often. Start with a smaller, cheaper model, measure quality on real examples, and move up only if you need to. Keep the model name in a config value or secret so you can change it without rewriting code. Check OpenAI's pricing page for current rates.
Options and trade-offs
| Decision | Simple choice | When to go further |
|---|---|---|
| Streaming | Return the whole answer | Stream for chat and long replies; users see text immediately |
| Conversation memory | Send recent messages each time | Summarise older messages when chats get long |
| Output format | Plain text | Ask for structured JSON when your code needs to read fields |
| Your own data | Include the relevant record in the request | Add search over your content when there's too much to include (see what RAG is) |
| Non-urgent bulk work | Loop over items | Use OpenAI's batch option for lower cost when results can wait |
Prompts to give your AI builder
A first feature:
Add a "Summarise" button on the ticket page. It calls a server route POST /api/summarise that checks the user is logged in, loads the ticket, and calls the OpenAI Responses API using the official openai SDK and the secret OPENAI_API_KEY. Read the model name from the secret OPENAI_MODEL. Instructions: "Summarise in 3 bullet points, then suggest a short reply. Don't invent order numbers." Cap the output length, stream the reply to the page, and never send the key to the browser.
Guardrails:
Limit each user to 30 AI requests per hour, stored in the database. Log input and output tokens per request with the user ID and feature name. If OpenAI returns 429, respect the Retry-After header and retry with exponential backoff and jitter up to 3 times; after that, show "The AI is busy, try again in a minute." Show a friendly error for any other failure and keep the rest of the page working.
Handling errors and limits
OpenAI measures rate limits in several ways, including requests per minute and tokens per minute, and you hit whichever runs out first. When you exceed one, the API returns a 429 with a Retry-After header. OpenAI's docs say to treat that value as a minimum and add a small random delay so clients don't all retry together.
Also plan for:
- Timeouts on long generations. Streaming helps, and so does a sensible output cap.
- Provider outages. The AI feature should fail gracefully without breaking the page.
- Bad output. Validate structured output before saving it, and don't let model output trigger sensitive actions without a check in your code.
Common mistakes
- Key in the frontend or in the repository. Browser code is public, and a leaked key can be abused within minutes. See how to keep API keys safe.
- No spend limit. A loop or an abusive user can run up a large bill quickly.
- No per-user limits. Your provider limit protects your wallet, not your other users.
- Trusting the prompt for security. Users can type "ignore previous instructions". Permissions belong in server code.
- Sending more than needed. Whole documents and long histories multiply cost. Send what the task needs.
- Hard-coding the model name. Models are updated and retired; make it a setting.
- Sending personal data without thinking. Read OpenAI's data-use terms and send only what the feature needs.
Checklist
- Separate OpenAI projects (or keys) for testing and production
- Key stored as a server-side secret, never in browser code or the repo
- Spend alerts and a hard spend limit set
- All calls go through your server and check the user's permissions
- Output length capped; model name configurable
- Per-user rate limit in your app
- 429 handling with Retry-After, backoff and jitter
- Friendly error state; the page works without the AI
- Token usage logged per feature
- Tested with long, empty and hostile inputs
Using the OpenAI API in Mythex
AI features in apps you build on Mythex use your own OpenAI key, billed to your OpenAI account rather than to Mythex build credits. Add OPENAI_API_KEY as a project secret through /secrets, project settings or the secrets card the agent shows when it needs a key, then describe the feature with a prompt like the ones above. After changing secrets on a published app, publish again so it picks them up. The docs recipe Add AI features with your own API keys has a starter prompt, and Secrets in chat explains the secrets card. Comparing providers? See how to integrate the Claude API.
Questions
Which OpenAI API should I use for a new app?
As of September 2026, OpenAI's quickstart presents the Responses API as the starting point for new projects, with examples using client.responses.create() in its official SDKs. Check OpenAI's documentation for the current recommendation before you start.
Can I call the OpenAI API from the browser?
Not with your secret key. Anyone can read browser code and reuse the key to run up your bill. The browser should call your own server route, and your server calls OpenAI with the key stored as an environment variable.
How do I stop the OpenAI API from costing too much?
OpenAI's production guide recommends spend alerts and a hard monthly spend limit on the limits page. In your app, cap reply length, limit requests per user, and log token usage so you can see which features cost the most.
What does a 429 error from OpenAI mean?
It means you hit a rate limit. OpenAI's rate-limit docs say the response includes a Retry-After header; wait at least that long, add a small random delay, and retry with exponential backoff.