TL;DR
Under Fluid compute, a function with maxDuration = 1800 starts a new instance roughly every 30 requests. The same code at 300s or 800s starts one every 600–700 requests. Requests routed to those new instances spend ~2.2–3.0s on the platform before the handler runs. This shows up as a separate latency peak at 2.25–3.25s, and it dominated our p95/p99. Moving everything except one genuinely long route to ≤800s removed it. Sharing the numbers in case they help others and the beta team.
Setup
- Next.js 16.3.5 App Router, Node 24, Fluid compute, 2048 MB, single region (bom1)
- One Hono app served by thin App Router route handlers. The same application code ran under four function layouts, so the only variable was
maxDurationgrouping (confirmed withvercel inspect):
| Period | Layout |
|---|---|
| A | Route handler per route; a few long routes at 1800s |
| B | All API routes in one function at maxDuration = 1800 |
| C | 7 functions; short routes on 300s; 4 entries still at 1800s |
| D | 4 functions: 300s / 790s / 800s / 1800s; only a low-traffic WebSocket voice proxy remains at 1800s |
How we measured
- Platform duration:
report.durationMsfrom the log drain. - Handler duration: timed inside our framework middleware, so it excludes anything before the handler runs.
- Instance identity: each process creates a UUID on
globalThiswhen our module first loads and logs{id, ageMs, rss}to stdout at most once per 30s. A new instance shows up as a new ID withageMsunder 5s. - Caveat: the drain's request ID can't be joined to our request events, so attribution is statistical (timing, distribution shape, instance age).
1. A single fixed ~2.5s mode across unrelated trivial routes
In layout B, five unrelated trivial routes (p50 of 12–100ms) all had p99 ≈ 2,783ms. In layouts A and D the same routes had p99 of 70–300ms.
Histogram for those routes in B:
| ms | 500 | 750 | 1000 | 1250 | 1500 | 1750 | 2000 | 2250 | 2500 | 2750 | 3000 | 3250 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| n | 20.4k | 12.0k | 3.9k | 523 | 152 | 117 | 1.6k | 10.0k | 26.7k | 15.8k | 7.3k | 3.8k |
It dips at 1.25–2s and has a separate peak at 2.25–3.25s. The handler-side p99 for one of these routes was 517ms while the platform reported 3,072ms. So the time is spent before our code runs.
2. The slow mode tracks new-instance creation
Hourly in layout B:
| New instances / hour | 2–4s requests / hour (short routes) |
|---|---|
| 46,137 | 10,180 |
| 34,475 | 6,816 |
| 21,328 | 4,376 |
| 12,534 | 2,697 |
| 1,217 | 132 |
- Instances served about one request each. 69% of instances served exactly one logged request. The median instance was last seen 2.5s after it started. Only 55% of requests reached an instance older than 5 minutes; in layout C it was 95%.
- No crashes. No exit, SIGKILL, out-of-memory or SIGTERM logs, and RSS was steady at ~216 MB.
- What the first request sees. When a new instance's first request reaches the handler, the process is already 1.2–2.2s old, so the request waited for boot.
3. Instance reuse by duration tier
Layout C, 24h:
| maxDuration | Requests | New instances | Requests per new instance |
|---|---|---|---|
| 300s | 6.77M | 10.1k | ~670 |
| 600s | 2.05M | 4.7k | ~430 |
| 1800s | 5.09M | 158k | ~32 (90% sampled only once) |
Layout D, 5h: 300s ≈ 700 requests per new instance, 800s ≈ 580. So 800s behaves normally. The drop-off is specific to the extended-duration beta.
4. Before/after moving short routes off 1800s (layout C → D, same hours, platform duration)
| Route group | p99 | share at 2–4s |
|---|---|---|
| CRUD GETs | 2,111ms → 239ms | 1.36% → 0.08% |
| CRUD writes | 1,990ms → 117ms | 0.98% → 0.06% |
New instances across the project in the same 5h window: ~50.5k → ~6.2k. Error rates unchanged.
Time to first SSE chunk for our streaming LLM endpoint, measured on the client:
| p50 | p90 | p95 | |
|---|---|---|---|
| Endpoint at 1800s | 286ms | 976ms | 4.3s |
| Endpoint at 300s | 269ms | 538ms | 1.17s |
Takeaways
- Don't put latency-sensitive or high-frequency routes in a function with
maxDuration > 800while it's in beta. Isolate only the routes that truly need it. - Watch for accidental grouping. Route handlers with matching config can be bundled into the same function. A catch-all or shared handler at 1800s puts everything it serves on the extended-duration path. Check with
vercel inspect <deployment> --json: comparelambda.functionNameandlambda.timeoutinbuilds[].output[]. - 800s (GA) reused instances normally in our data.
Questions for the team
- Is the aggressive instance cycling for >800s functions intended during the beta? Will it change before GA?
- Are these instances terminated after invocations, or kept alive but not routed to?
- Is there a way to get normal instance reuse for a function that needs 1800s (a WebSocket proxy for voice calls up to ~30 min)?
Happy to share deployment and request IDs privately with Vercel staff.