In short. A client can measure how long the server took but never what it was doing. Server-Timing is a response header that carries a breakdown of that time — database, cache, template, upstream call — to the browser, where DevTools displays it and JavaScript can read it. It turns the largest interval in a page load from a single number into an account.
The one interval no external measurement can open
Timing a request from outside gives you DNS, connect, TLS and then a gap before the first byte. That gap is the server thinking, and from the client it is opaque: you know it was 600 ms and nothing about why.
Server logs can answer it, but only for whoever has the logs, only after the fact, and never for the specific slow request a user just complained about. Server-Timing closes that by having the server annotate its own response.
This is what makes it different from ordinary logging. The breakdown travels with the response, so it is available in the same place as every other client-side metric — attached to the exact request, in the browser of the person who experienced it.

What the header looks like
A comma-separated list of metrics. Each has a name, an optional duration in milliseconds, and an optional description.
Server-Timing: db;dur=53, tpl;dur=12.4, upstream;dur=210, cache;desc="MISS"
# name only — useful for flags rather than durations
Server-Timing: cache;desc="HIT"
# durations are milliseconds, fractional values are allowed
Server-Timing: total;dur=284.7
Nothing about the names is standardised — they are yours. That freedom is the point and also the trap: names chosen ad hoc across services produce a breakdown nobody can compare. Agree on a small vocabulary and reuse it everywhere.
Where it becomes visible
| Surface | What you get | Requires |
|---|---|---|
| Chrome DevTools, Network → Timing | The metrics rendered next to the network phases | Nothing |
PerformanceServerTiming in JS | Programmatic access for your own RUM | Same origin, or Timing-Allow-Origin |
| Monitoring and RUM products | Aggregation across real visits | Product support |
// Read the breakdown for the navigation itself
const nav = performance.getEntriesByType('navigation')[0];
for (const m of nav.serverTiming) {
console.log(m.name, m.duration, m.description);
}
// And for a specific subresource
for (const r of performance.getEntriesByType('resource')) {
if (r.serverTiming.length) console.log(r.name, r.serverTiming);
}
The cross-origin rule catches people out. DevTools shows the header for any response, but JavaScript can only read
serverTimingfor a cross-origin resource when that response also carriesTiming-Allow-Origin. If your API is on another host, your own RUM sees nothing until you add it there.
What is worth measuring
The useful breakdown answers "which part of the work dominated", not "how long did everything take" — the total is already measurable from outside. Four or five metrics is usually enough:
- Database — total query time for the request, not per query.
- Cache — hit or miss as a description, which explains a slow request better than any duration.
- Upstream — time spent waiting on another service. This is the one that most often turns out to be the answer.
- Render — template or serialisation cost.
- Queue — time the request waited before a worker picked it up, if your stack can see it.
The last one is worth adding wherever possible. A request that spent 400 ms queued and 40 ms working looks identical from outside to one that worked for 440 ms, and the two have completely different fixes — more capacity against faster code.
What you are exposing, and to whom
The header goes to every client, including ones you did not have in mind. Metric names leak architecture: elasticsearch, legacy-billing-api and redis-session each tell a reader something about your stack, and durations reveal which dependencies are slow enough to be worth probing.
Two reasonable postures, and the choice depends on how much your surface is worth attacking:
# nginx — emit the detailed breakdown only for authenticated staff
map $cookie_debug_timing $timing_detail {
default "";
"1" $upstream_response_time;
}
add_header Server-Timing "app;dur=$timing_detail" always;
# Or emit an opaque, useful-but-uninformative set to everyone
Server-Timing: a;dur=53, b;dur=210, c;desc="MISS"
Neither is wrong. What is wrong is publishing internal service names by accident because the header was added in development and never reviewed before it shipped.

Adding it to a real stack
nginx
# The proxy already knows how long the upstream took
add_header Server-Timing "upstream;dur=$upstream_response_time000" always;
# $upstream_response_time is in seconds — the trailing zeros convert to ms
PHP
<?php
$t0 = microtime(true);
// ... query the database ...
$db = (microtime(true) - $t0) * 1000;
$t1 = microtime(true);
// ... render ...
$tpl = (microtime(true) - $t1) * 1000;
header(sprintf('Server-Timing: db;dur=%.1f, tpl;dur=%.1f, cache;desc="%s"',
$db, $tpl, $wasCached ? 'HIT' : 'MISS'));
Node
const t0 = process.hrtime.bigint();
// ... work ...
const ms = Number(process.hrtime.bigint() - t0) / 1e6;
res.setHeader('Server-Timing', `app;dur=${ms.toFixed(1)}`);
Set the header before the response body starts, since headers cannot be added once output has begun. In streamed responses that means measuring the work that happens before the first byte and accepting that anything after it is not represented.
Using it to finish a TTFB investigation
Measuring from outside tells you the server-side interval dominates. Server-Timing tells you what inside it dominates, and together they close the question:
# Outside: how much of the total is the server?
curl -o /dev/null -s -w 'ttfb:%{time_starttransfer}s total:%{time_total}s\n' https://example.com/
# Inside: what did the server spend it on?
curl -sSI https://example.com/ | grep -i '^server-timing:'
The pairing is what makes both useful. A large server interval with cache;desc="MISS" and a large upstream duration is a different problem from the same interval with a hit and a large database time — the first is a dependency, the second is your own query. The stage-by-stage view from outside is covered in the TTFB breakdown; this header is how you open its last stage.

Reading the breakdown: four patterns and what each one means
| What the header shows | What it means | Where the fix is |
|---|---|---|
| Large upstream, small everything else | You are waiting on another service | That service, or removing it from the response path |
| Large database, cache miss | The query is genuinely expensive and nothing absorbed it | Indexing, or caching the result |
| Large database, cache hit | The cache is being consulted and the query still runs | The caching logic — a hit that does not short-circuit |
| Large queue, small work | Requests wait before a worker takes them | Capacity or concurrency, not code |
| Everything small, interval still large | Time is being spent where you are not measuring | Add a metric around the unmeasured span |
The third row is the one worth looking for deliberately. A cache that reports a hit while the database time stays high means the lookup happens but the result is not used — an expensive bug that looks like a working cache from every angle except this one.
The last row is the honest reminder that a breakdown only describes what you instrumented. If the named metrics sum to far less than the observed server interval, the missing time is real and unmeasured, and the next metric to add is the one that covers the gap.
What your stack may already be sending
Before instrumenting anything, look at what arrives today. Several CDNs, application frameworks and hosting platforms emit their own metrics without being asked — cache status, edge processing time, origin wait — and reading them costs one request.
# What does your site send right now?
curl -sSI https://example.com/ | grep -i '^server-timing:'
# Compare origin against edge: the two often carry different metrics
curl -sSI --resolve example.com:443:203.0.113.10 https://example.com/ | grep -i '^server-timing:'
Where an edge already reports cache status, adding your own duplicate is noise. Fill the gaps instead: the edge can describe what it did, but only your application can describe what happened inside it.
How to check your site right now
Use the HTTP header checker to see whether your responses already carry Server-Timing and what it says — many frameworks and CDNs emit one you did not add, and reading it is the fastest way to find out what your stack already knows about itself. The speed check gives the outside timing breakdown that the header is meant to complement.
The header describes one request at a time, which makes it excellent for diagnosis and useless for trends. Scheduled monitoring records response time continuously, so you can tell whether the slow request you just inspected is the normal state or an outlier — a distinction no single measurement can make.

Frequently asked questions
Does Server-Timing slow anything down?
The header itself is a few dozen bytes. What can cost is the instrumentation behind it — timing every individual query rather than the aggregate. Measure at the level you will act on, which is usually four or five totals rather than a trace.
Why can DevTools see it when my JavaScript cannot?
Cross-origin protection. DevTools displays the raw response, but PerformanceServerTiming is only exposed to script for cross-origin resources that also send Timing-Allow-Origin. Add that header on the other host and it appears.
Is it safe to expose in production?
The durations are harmless; the names can describe your architecture. Either use opaque labels for everyone or emit the detailed set only for authenticated staff. The failure to avoid is shipping development labels unreviewed.
Are the metric names standardised?
No, and that is intentional. The cost is that inconsistent names across services produce a breakdown you cannot compare, so agree on a short shared vocabulary before instrumenting the second service.
Can I add it after the response has started?
No — headers are sent before the body. For streamed responses you can only report work completed before the first byte, which is usually what you wanted to measure anyway.
Do CDNs strip it?
Some pass it through, some add their own metrics, some drop unknown headers. Check what actually arrives at the browser rather than what your origin sends, because the two can differ.
Checklist
- Measure aggregates, not every individual query.
- Include a cache hit or miss description — it explains more than most durations.
- Add queue time wherever the stack can see it.
- Agree on metric names before instrumenting a second service.
- Add
Timing-Allow-Originif your own scripts need to read cross-origin timings. - Review the names before they ship — they describe your stack to everyone.
- Set the header before output begins.
- Check what arrives at the browser, not what the origin sends.
- Pair it with an outside measurement; neither is complete alone.
- Use a recorded series to tell an outlier from the normal state.