Skip to content
SELFOIA
Free check
Open menu

Speed & AI-Cost Sprint · from $3,000 · 1–2 weeks

Make your AI‑built app fast and your AI bill predictable.

The Speed & AI-Cost Sprint is for founders whose AI bill keeps climbing or whose app slows down as users arrive. We put limits on your AI features, cut wasted tokens (the units AI providers bill you by), and fix slow pages and database queries. From $3,000 with a fixed estimate, done in 1–2 weeks, with before-and-after numbers.

We measure before we change anything · Fixed estimate · Tested on a copy, not your live app

Sound familiar?

  • “My AI bill doubled and I don't know why.”

    What's usually behind it

    Usually there's no limit on who can call your AI, or a key leaked. We rule out a leak first.

  • “Someone could use my AI for free and I'd pay the bill.”

    What's usually behind it

    If your AI feature has no cap per user, one person or a bot can run it all night.

  • “My AI costs grow faster than revenue.”

    What's usually behind it

    Each request may be sending the whole chat history, retrying on errors, or using the most expensive model for a simple job.

  • “It gets slow when more people use it.”

    What's usually behind it

    Missing indexes, pages that load everything, no cache.

  • “It's fast for me, slow with real data.”

    What's usually behind it

    Your test account has a little data; real accounts have a lot. Lists that load every row get slower as your data grows.

  • “The site is slow on phones.”

    What's usually behind it

    Heavy code and oversized images. Google checks this with Core Web Vitals, its speed and stability checks for real visitors.

The bill and the wait are both real.

$600,000

A login bug in one vibe-coded internal dashboard let an attacker steal an AI key and use about $600,000 worth of donated AI credits over three weeks.METR, 2026 (source, opens in a new tab)

84%

In a 2025 survey of 372 mostly software companies, 84% said AI costs cut their gross margins by more than 6 points.Mavvrik & Benchmarkit, 2025 (source, opens in a new tab)

48%

Only 48% of websites pass Google's Core Web Vitals speed and stability checks for visitors on phones, so more than half do not.HTTP Archive Web Almanac, 2025 (source, opens in a new tab)

Every number on this site links to its source.

What we do, zone by zone

We start with the money, because a runaway bill can't wait. Then we make the app fast for the users you already have.

Money: your AI bill

  • E01Rule out a leaked key first. If your bill jumped overnight, we check whether your key is visible in your app or your public code, and help you replace it.
  • E11Put limits on your AI features: log-in required, a cap per user per hour and per day, and a maximum length for each request.
  • E11Set a budget alert at your AI provider, so you hear about a spike from an alert, not from your bill.
  • E13Measure what each user and each feature costs you.
  • E13Cut waste: send less chat history with each request, stop retry loops, save and reuse answers to repeated questions (caching), and use a smaller, cheaper model for simple jobs like tagging or short summaries.
  • E12Keep your chatbot from being talked into things: no secrets in its instructions, and access only to the data it needs.

Speed: pages and database

  • E14Add missing indexes. Like the index at the back of a book, they let your database find rows without reading every one.
  • E14Load long lists page by page, filter in the database instead of the browser, and stop the app from asking the database the same thing over and over.
  • E15Make pages lighter on phones: less code to download, right-sized images, a faster first view. We measure it with Core Web Vitals.
  • E16Make sure Google and AI search can read your main content.
  • E17Remove free-plan limits that bite on launch day: paused projects, sign-up email limits, no email sender of your own.

Safety net

  • E20Alerts when AI spend or errors jump, so you hear it from an alert, not a customer.

See all 24 checks

What we measure, before and after

We measure your app before we change anything, then again after, on the same pages and the same kinds of requests. You get both sets of numbers in your report. We don't promise a result before we've measured.

  • Mobile LCP (Largest Contentful Paint)

    How long your main content takes to appear on a phone. It's one of Google's Core Web Vitals.

  • p95 request time

    The time within which 95 out of every 100 requests to your app finish. It shows the slow end your users actually feel, not the average.

  • $ per 1,000 AI requests

    What your AI provider charges for every 1,000 AI requests your app makes. It shows whether each request got cheaper, not just whether traffic dropped.

Where your data allows, we also show what one active user costs you in AI per month.

We measure on a staging copy (a private copy of your app with fake data) and with reports you can export from your live app. We never load-test your live app.

What you get

  • The changes, in your own code: limits, caching, model changes, database and page fixes.

  • A before/after report with the three numbers above. Every report is read and signed by Denys.

  • A cost map: which features and which users cost you the most in AI.

  • Alerts in your own accounts for budget spikes and errors.

  • A short watch list: what to keep an eye on as you grow.

How a Speed & AI‑Cost Sprint works

  1. 1

    Tell us what's slow or costly.

    Start with a free check. A screenshot of your AI provider's usage page helps. We never need the key itself.

  2. 2

    We measure and quote.

    With read-only access, we measure your starting numbers and send a fixed estimate before any work starts.

  3. 3

    We fix and measure again.

    In 1–2 weeks you get the changes and the before/after numbers.

Your audit fee counts toward a Fix Sprint or Speed & AI-Cost Sprint booked within 30 days.

Price

Speed & AI-Cost Sprint

from $3,000

fixed estimate

1–2 weeks

You get the price in writing before we start, and it doesn't change unless you add work. If we find something bigger along the way, we tell you and quote it separately.

Paid by Stripe invoice once you approve the fixed estimate.

Fix Sprint · On-call CTO · Compare all prices

Included

  • AI limits, budget alerts and cost measurement
  • Caching, model choices and less wasted tokens
  • Database and page speed fixes
  • Before/after report with your own numbers, signed by Denys

Not included

  • Your AI provider, database or hosting bills
  • Security fixes outside speed and cost (that's a Fix Sprint)
  • Rebuilding your app on a new stack
  • New features
  • Watching your bills after the sprint (that's On-call CTO)
  • Load tests on your live app
  • A promised saving before we've measured

Who does the work

Denys Kharkovskyy · Founder & lead engineer

SELFOIA is led by Denys Kharkovskyy, founder & lead engineer. In the 20+ AI-built apps we've reviewed, the same problems keep coming back: AI bills that climb, pages that slow down, logins that break. Every report is read and signed by Denys.
  • 10+ years building software
  • 20+ AI-built apps reviewed
  • Worked on a real-time analytics platform used by 200+ companies
  • 5.0 on Upwork (opens in a new tab)
  • A team lead who reviews other engineers' code every day
About Denys

Questions about speed and AI costs

How do I stop strangers running up my AI bill?

Require log-in for every AI feature, cap requests per user per hour and per day, limit how long each request can be, and set a budget alert at your AI provider. That's what OWASP recommends for “unbounded consumption”, and it's the first thing we do in a sprint. Supabase's built-in limits protect sign-ups and logins; a cap on each user's AI requests is something your app has to add itself.

Why did my AI bill spike?

The usual suspects are a leaked key someone else is using, an AI feature with no limit per user, or a change that made each request bigger, like sending the whole chat history every time. We check them in that order, starting with the leak.

Can my app handle 1,000 users?

We can't answer that honestly without looking, because it depends on what each user does. What we can do is find what breaks first, fix it, and measure before and after on a staging copy with realistic fake data.

How do you measure the improvement?

With three numbers, before and after: mobile LCP (how fast your main content appears on a phone), p95 request time (how long the slowest 5 in 100 requests take), and dollars per 1,000 AI requests. You get both sets in your report.

Will a cheaper AI model make my app worse?

Not if the job fits. We only move a task to a cheaper model after comparing the answers side by side on your own examples, and you approve each switch. Hard tasks stay on the stronger model.

Do you need my OpenAI or Anthropic key?

No. We read your code and your usage reports; a screenshot or export from your provider's dashboard is enough. If you send us a key by mistake, we'll ask you to replace it.

Can you promise I'll save a certain percentage?

No. We don't promise numbers before we've measured. Your estimate lists what we'll change, and your report shows what actually changed, in your own numbers.

What happens after the sprint?

The limits and alerts keep working on their own. If you want someone to watch your bills and errors every month and review changes before they ship, that's On-call CTO.

Bill climbing? App slowing down? Start with a free check.

Send us your app. Within 2 business days you get a written reply with the biggest risks we can see from outside, including anything that could be running up your bill. Free, no call needed.