24 min readpublished

The First Token Is the Hardest

Cold starts, runtime stability, and the physics of serving large models

We spent twenty years making web servers boring. Stateless, small, fast to boot, and cheap to leave idle. Large models break every one of those assumptions at once, and most of the instability people blame on "AI" is really physics we haven't budgeted for.

AI InfrastructureInferenceSystems DesignMath

Here's a scene most platform teams have lived through. An endpoint has been quiet all night. At 8:02 a.m. someone opens the product, types a question, and watches a spinner for three minutes. By 8:05 the dashboard is fine, the on-call engineer sees nothing wrong, and the user has already left. Nothing crashed. The system did exactly what we built it to do. It just took three minutes to become a system.

A container running a web API is typically tens of megabytes and boots in well under a second. A container serving a 70-billion-parameter model has to pull an image that is often over ten gigabytes, initialize eight GPUs and the links between them, move about 140 GB of weights into high-bandwidth memory, compile kernels, capture execution graphs, and warm up before it can emit a single token. That is a gap of four to five orders of magnitude in state that has to exist before the first byte of response. Every pattern we inherited from stateless compute, like scale-to-zero, aggressive bin packing, and "just add a replica," was tuned for the small end of that gap.

This piece is my attempt to lay the whole problem out flat. I don't think "cold starts are slow" is one problem. I count at least eight dimensions, each with its own unit of measure, and they compound. Fixing one while ignoring the others is how teams end up with a fast boot and a system that still falls over at 9 a.m.

The eight dimensions of model runtime stability
DimensionThe questionUnit
TimeHow long until a new replica can answer?seconds per cold start
BytesHow much has to move, and over which link?GB ÷ GB/s
TrafficHow often does a request find nobody home?P(cold)
LoadWhat happens to latency as warm replicas fill up?1 / (1 − ρ)
MemoryHow many conversations fit at once?KV bytes per token
NeighborsHow much does everyone else slow me down?tokens/s per stream
NumericsDo I get the same answer twice?reduction order
FleetHow often does a user feel the worst case?1 − pⁿ

1. Anatomy of a cold start

A "cold start" is really seven separate waits stacked end to end. Some are about getting a machine, some are about moving bytes, and some are about rebuilding runtime state that was identical the last time this model booted:

Tcold=Tprovision+Timage+Tinit+Tfetch+Tload+Tcompile+TwarmupT_{\text{cold}} = T_{\text{provision}} + T_{\text{image}} + T_{\text{init}} + T_{\text{fetch}} + T_{\text{load}} + T_{\text{compile}} + T_{\text{warmup}}

Figure 1 models that sum for four increasingly prepared deployments. The worst case, a brand-new GPU node pulling everything from object storage, lands around seven minutes. The best case, with weights already pinned in host memory and compiled graphs restored from cache, is about twenty seconds. That is a 19× difference on identical GPUs running an identical model.

0s60s120s180s240s300s360s420sA new GPU node, weights in object storage420sB warm node, weights in object storage180sC warm node, weights on local NVMe, compile cache73sD weights pinned in host RAM, graphs restored22s
node provision
image pull
runtime + NCCL init
weight fetch
deserialize → HBM
compile + CUDA graphs
warmup
Fig. 1 Modeled cold-start budget for Llama 3 70B (bf16, 140 GB) on one 8× H100 node. Blue stages move bytes, dark stages build the runtime around them. Stage times are derived from spec-sheet bandwidths and typical serving-stack behavior, not a benchmark of any vendor. The key point is the shape: every tier removes a whole category of work, not a few percent.

Two things stand out to me. First, the GPU is idle for almost the entire timeline. Every bar except the last one is disk, network, CPU, or a scheduler. You're renting the most expensive silicon in the building so it can wait on I/O.

Second, the dark "compile + CUDA graphs" segment is recomputation. Serving engines capture execution graphs for a set of batch sizes so each decode step skips per-kernel launch overhead, and many stacks also run a compiler pass. That work depends only on the model, the GPU type, and the parallelism layout, and none of those changed since the last boot. When a stage produces the same output every time, it belongs in a cache, not on the critical path.

2. Bytes over bandwidth

Strip away orchestration and the fetch and load stages reduce to one line of physics:

Ttransfer≥Nparams×bBbottleneckT_{\text{transfer}} \geq \frac{N_{\text{params}} \times b}{B_{\text{bottleneck}}}

Here bb is bytes per parameter (2 for bf16, 1 for fp8) and BbottleneckB_{\text{bottleneck}} is the slowest link on the path. For our reference model, 70.6×109×2≈14170.6 \times 10^9 \times 2 \approx 141 GB. The table of possible links spans more than two orders of magnitude:

∼ ⁣1 GBs⏟one object-store stream  ≪  ∼ ⁣7 GBs⏟NVMe Gen4  ≪  ∼ ⁣25 GBs⏟PCIe Gen4 ×16  ≪  3,350 GBs⏟H100 HBM3\underbrace{\sim\!1\ \tfrac{\text{GB}}{\text{s}}}_{\text{one object-store stream}} \;\ll\; \underbrace{\sim\!7\ \tfrac{\text{GB}}{\text{s}}}_{\text{NVMe Gen4}} \;\ll\; \underbrace{\sim\!25\ \tfrac{\text{GB}}{\text{s}}}_{\text{PCIe Gen4}\ \times16} \;\ll\; \underbrace{3{,}350\ \tfrac{\text{GB}}{\text{s}}}_{\text{H100 HBM3}}

Over a single object-store stream the model takes about 140 seconds to arrive. Over local NVMe it takes about 20. From pinned host memory each GPU only needs its own 1/8 shard, about 17.5 GB, over its own PCIe link, which is well under a second in theory. In practice, deserialization and allocation overhead push that into single-digit seconds. That is also why a zero-copy, memory-mappable format like safetensors [13] matters more than it looks.

0501001502000100200300400checkpoint size (GB)seconds8B bf1670B bf16405B fp8
object store, single stream (~1 GB/s)
object store, parallel ranged GETs (~5 GB/s)
local NVMe Gen4 (~7 GB/s)
host RAM → GPU, PCIe Gen4 ×16 (~25 GB/s)
Fig. 2 Pure transfer time, t = bytes ÷ bandwidth, for three common checkpoint sizes. This is the floor before deserialization, sharding, or compilation. A single-stream download of a 70B checkpoint alone takes over two minutes, and a 405B checkpoint leaves the chart. Storage tier is a first-order design decision, not an implementation detail.

The uncomfortable implication is that model size grows cold start linearly unless you change tiers. A team that moves from an 8B to a 405B model on the same infrastructure has quietly multiplied its transfer floor by 25. ServerlessLLM [3]attacks exactly this. It treats the GPU server's local memory and storage hierarchy as a checkpoint cache, uses a loading-optimized format, and schedules requests toward the servers that already hold the bytes. The general lesson is locality-aware scheduling: route the request to where the weights already are, rather than moving weights to where the request landed.

3. How often you pay: the traffic dimension

A cold start only hurts when a request actually waits on one. Suppose a replica scales to zero after an idle timeout TT and requests arrive as a Poisson process at rate λ\lambda. The gap between arrivals is exponentially distributed, so the chance a request arrives after the replica has gone cold is

P(cold)=P(gap>T)=e−λTP(\text{cold}) = P(\text{gap} > T) = e^{-\lambda T}
00.250.50.751051015202530idle timeout T (minutes)P(cold)13.5% cold at T = 10 min
λ = 0.05 / min (one call every 20 min)
λ = 0.2 / min (every 5 min)
λ = 1 / min
Fig. 3 Probability that a request lands on a scaled-to-zero replica, P(cold) = e^(−λT), for Poisson arrivals at rate λ and an idle timeout T. Low-traffic endpoints stay cold most of the time even with generous timeouts. Real traffic is burstier than Poisson, which makes the long quiet gaps longer and the curve worse.

An endpoint that sees one request every five minutes, with a ten-minute timeout, sends e−2≈13.5%e^{-2} \approx 13.5\% of its users into a cold start. One request every twenty minutes means e−0.5≈61%e^{-0.5} \approx 61\%. That long tail of quiet endpoints is not hypothetical. In Microsoft's characterization of production serverless traffic, most applications were invoked far less than once per minute on average, and invocation rates varied by many orders of magnitude between applications [4]. Model endpoints are typically lumpier still: internal tools, per-customer fine-tunes, evaluation jobs, and features that only wake up during business hours.

There's a subtler result hiding in that formula that I think most autoscaling policies miss. The rate of cold starts, not the probability, is what costs you GPU time and pages the on-call:

Rcold(λ)=λ e−λT,dRdλ=e−λT(1−λT)=0  ⇒  λ∗=1T,Rmax⁡=1eTR_{\text{cold}}(\lambda) = \lambda\, e^{-\lambda T}, \qquad \frac{dR}{d\lambda} = e^{-\lambda T}(1 - \lambda T) = 0 \;\Rightarrow\; \lambda^{*} = \frac{1}{T}, \quad R_{\max} = \frac{1}{eT}

The endpoints that generate the most cold starts are neither your busiest nor your quietest. They are the ones whose traffic rate is about one request per idle timeout. With a 10-minute timeout, an endpoint at roughly six requests an hour can churn through more than two cold starts every hour, forever. The busy endpoint never goes cold, and the dead one rarely wakes. The middle is where your fleet flaps.

4. Spikes: where cold start becomes a queue

Scale-from-zero is the visible case. The more damaging case is scale-from-one during a traffic spike, because every second of cold start becomes queued requests. Let traffic jump to rate λ\lambda against one warm replica with service rate μ<λ\mu < \lambda. The autoscaler reacts immediately and requests a second replica, which becomes useful after TcT_c. The backlog is a triangle:

Qpeak=(λ−μ) Tc,Tdrain=(λ−μ) Tc2μ−λQ_{\text{peak}} = (\lambda - \mu)\, T_c, \qquad T_{\text{drain}} = \frac{(\lambda - \mu)\, T_c}{2\mu - \lambda}

With λ=6\lambda = 6 and μ=4\mu = 4requests per second, a 30-second cold start peaks at 60 queued requests and is fully recovered a minute after the spike. A 180-second cold start peaks at 360 queued requests and takes six minutes to recover. By Little's law L=λWL = \lambda W [5], the average wait grows in proportion to the queue.

0100200300400060120180240300360seconds since spikequeued requestspeak 60peak 360
cold start = 30 s
cold start = 180 s
Fig. 4 Queue depth when traffic jumps from 2 to 6 req/s against one warm replica serving 4 req/s. The autoscaler adds a second replica immediately, but it only serves after its cold start. The backlog grows at λ − μ until the replica is ready and drains at 2μ − λ afterward. Peak depth and recovery time both scale linearly with cold-start time: six times the cold start means six times the pain, twice.

The linear model is optimistic, because real clients don't wait politely. If client timeouts are shorter than the cold start, users and SDKs retry. Each retry is a new arrival, so the effective λ\lambdarises exactly when capacity is lowest, and the triangle turns into a ramp. This is the mechanism behind most "the model provider was down" incidents I've seen: a slow boot, short timeouts, and retry logic combine into a retry storm. The fix is less about faster GPUs and more about making cold start shorter than your client timeout and giving clients a clear signal to back off.

5. Warm isn't the same as stable

Say you solve cold starts completely and every replica is always warm. The next instability is the one finance asks for. GPUs are expensive, so the pressure is to run them hot. Queueing theory has been unambiguous about the cost of that since the 1970s [6]. For the simplest model, a single server with random arrivals and service times (M/M/1), the mean time in the system is

W=S1−ρ,ρ=λμW = \frac{S}{1 - \rho}, \qquad \rho = \frac{\lambda}{\mu}

where SS is the service time and ρ\rhois utilization. At 50% utilization, requests spend twice their service time in the system. At 80% it's 5×, at 90% it's 10×, and at 95% it's 20×.

0510203000.20.40.60.81utilization ρlatency × service timethe cliff2×5×10×20×
Fig. 5 Mean time in system for an M/M/1 queue, in multiples of service time: W / S = 1 / (1 − ρ). Latency roughly doubles from 50% to 80% utilization and doubles again from 80% to 90%. Squeezing the last 10% of GPU utilization out of a warm fleet costs about as much latency as the first 90% combined.

LLM serving is kinder than M/M/1 in one way and crueler in another. Continuous batching lets one replica serve many requests at once, which raises effective μ\mu. But service time varies enormously, since one request asks for 20 tokens and the next asks for 4,000. Variance is what feeds a queue. In the Pollaczek–Khinchine form for general service times, mean waiting time scales with (1+Cs2)/2(1 + C_s^2)/2, where CsC_s is the coefficient of variation of service time:

Wq=ρ1−ρ⋅1+Cs22⋅SW_q = \frac{\rho}{1 - \rho} \cdot \frac{1 + C_s^2}{2} \cdot S

Output lengths with a coefficient of variation of 2 make queueing delay 2.5× worse than the same average load with fixed-length work. So the right utilization target for a model fleet is lower than it is for a web fleet, not higher. If your capacity plan says "run GPUs at 90% to justify the spend," your latency plan has already been decided for you.

6. Memory: the KV cache is the real capacity

GPU memory holds two things: weights, which are fixed, and the key/value cache, which grows with every token of every active conversation. The per-token cost comes straight from the architecture:

KV bytes/token=2×nlayers×nkv heads×dhead×b=2×80×8×128×2=327,680≈320 KiB\text{KV bytes/token} = 2 \times n_{\text{layers}} \times n_{\text{kv heads}} \times d_{\text{head}} \times b = 2 \times 80 \times 8 \times 128 \times 2 = 327{,}680 \approx 320\ \text{KiB}

The leading 2 is for keys and values. Grouped-query attention, with 8 KV heads instead of 64, already makes this 8× smaller than it would otherwise be [8]. Our node has 8×80=6408 \times 80 = 640 GB of HBM. Reserve about 10% for headroom, subtract 140 GB of weights and some activation and graph workspace, and roughly 400 GB is left for KV. That buys about 1.2 million tokens of live context, shared by everyone on the node.

02004006005962k1498k3732k9128kcontext length per sequence (tokens)
Fig. 6 Maximum sequences that fit in a ~400 GB KV-cache budget at full context, for Llama 3 70B (80 layers, 8 KV heads, head dim 128, bf16 = 320 KiB per token). The node has the same GPUs, weights, and price in every column; only the context length changes. Going from 2k to 128k divides concurrency by 64.

That's why context length is a capacity decision, not a product checkbox. The same hardware serves 596 concurrent 2k-token chats or 9 concurrent 128k-token document sessions. When demand exceeds what fits, the engine has to preempt a sequence, either swapping its cache out or throwing it away and recomputing it later. The user sees a stall with no error. PagedAttention [2]was a major step because it removed the fragmentation that used to waste much of this memory, so more of the budget holds real tokens. It can't change the arithmetic of the budget itself.

7. Your latency depends on your neighbors

Web requests on the same server are mostly independent. LLM requests in the same batch are not. Decoding one token for every sequence in the batch requires reading all the weights from HBM once, plus each sequence's entire KV cache. At small and medium batch sizes that step is limited by memory bandwidth, not compute [9], so the time per step and each stream's speed are roughly

tstep≈W+∑i=1Bci⋅kBHBM,vstream=1tstept_{\text{step}} \approx \frac{W + \sum_{i=1}^{B} c_i \cdot k}{B_{\text{HBM}}}, \qquad v_{\text{stream}} = \frac{1}{t_{\text{step}}}

Here WW is weight bytes, cic_i is the context length of sequence ii, and kk is KV bytes per token. Alone on the node with 4k of context, a stream tops out near 26.8 TB/s÷141 GB≈19026.8\,\text{TB/s} \div 141\,\text{GB} \approx 190tokens per second. Put 64 such streams in the batch and each drops to about 119 tokens per second while aggregate throughput rises to about 7,600. At 256 streams, each gets about 55. A quick compute check confirms this is still bandwidth-bound: 256 tokens × ~140 GFLOP ≈ 36 TFLOP per step, about 4.5 ms of the node's dense bf16 peak, versus about 18 ms spent reading memory.

05010015020013264128192256concurrent sequences in the batchtokens / s per stream~190 tok/s alone
1k tokens of context each
4k tokens each
16k tokens each
Fig. 7 Bandwidth-bound ceiling on per-stream decode speed for Llama 3 70B on 8× H100 (26.8 TB/s aggregate HBM), as the batch fills. Every decode step reads all the weights once plus every active sequence's KV cache. Your tokens per second depend on how many other people are in the batch and how long their conversations are. The 16k line stops where the KV budget runs out.

This is the property I find most underappreciated. A user's tokens-per-second depends on how many strangers share their batch and how long their conversations are.The same prompt at the same hour can stream at 190 or 55 tokens per second, and nothing is broken either way. Operators trade per-user speed for aggregate throughput every time they raise the batch limit. That's a legitimate trade, but it should be a stated SLO, not an accident.

Prefill makes it worse. Processing a new 32k-token prompt is a large, compute-bound burst. If the scheduler puts it into the same iteration as a hundred decoding streams, every one of those streams stalls for that iteration, and users watching output see the text freeze mid-sentence. Sarathi-Serve [10] splits prefills into chunks that fit alongside decodes. DistServe [11]moves prefill and decode onto separate GPU pools and optimizes goodput, meaning requests that meet their latency target, rather than raw throughput. Both start from the same observation: the two phases have opposite hardware profiles, and mixing them naively turns one user's long prompt into everyone else's latency spike.

8. Numerics: same prompt, different answer

Stability isn't only about time. Teams are routinely surprised that the same prompt at temperature zero can produce different outputs on different calls. The root cause is that floating-point addition isn't associative:

(1020+(−1020))+1=1but1020+((−1020)+1)=0(10^{20} + (-10^{20})) + 1 = 1 \qquad\text{but}\qquad 10^{20} + ((-10^{20}) + 1) = 0

GPU kernels sum long vectors in parallel, and the order of those sums can depend on how the work is split. The common assumption is that concurrency randomness in atomics is to blame. He and colleagues at Thinking Machines argue that the dominant cause in typical inference engines is different: kernels that aren't batch-invariant, where the reduction strategy for your request changes with batch size [12]. Batch size changes with server load, which you don't control. So your output depends, in the last few bits, on who else is using the service. Greedy decoding amplifies last-bit differences: once two logits swap order, the sequences diverge and never reconverge.

Across a model upgrade, or a change in tensor-parallel layout, or a new GPU type, the differences get larger. That's why evaluation results don't transfer cleanly across deployments that should be "the same model." Batch-invariant kernels exist and cost some throughput. I'd treat them like strong consistency in a database: something you enable on the paths that need reproducibility (evals, audits, caching, regression tests), not a global default you either always pay for or never get.

9. The fleet: tail at scale, now with agents

Dean and Barroso's observation from 2013 is more relevant to AI systems than to the search backends it was written about [1]. If each call is slow with probability 1−p1 - p, a request that makes nn calls hits at least one slow call with probability

P(slow request)=1−p nP(\text{slow request}) = 1 - p^{\,n}
00.250.50.75111020406080100model calls per user request (n)P(≥1 slow)63% at n = 10018% for a 20-step agent
each call slow 1% of the time (p99)
each call slow 0.1% of the time (p99.9)
Fig. 8 Probability that a request touching n model calls hits at least one slow call, 1 − pⁿ. The math is the same whether the calls fan out in parallel (retrieval, ensembles, guardrails) or run in sequence (agent loops). At 100 calls, a 1-in-100 tail becomes the typical experience.

In 2013, nnwas the fan-out of a search query across index shards. In 2026, it's the number of model calls behind one user action: a retrieval step, a reranker, a guardrail, a planner, and a dozen tool-calling turns in an agent loop. A 20-step agent whose every call meets a p99 target still gives 18% of users a p99 experience somewhere in the run. At 100 calls, 63% do. In a sequential agent, the cost is worse than in a parallel fan-out because slow steps add up instead of overlapping. Every dimension above, from cold starts and queueing to KV preemption and noisy neighbors, feeds that 1−p1 - p, and agents raise it to a power.

What to actually measure

Most model dashboards I review show GPU utilization, average latency, and error rate. By the argument above, all three can look healthy while users suffer. These are the numbers I'd put on the first screen:

Metrics for model runtime stability
MetricDefinitionWhy it matters
Cold start, by stageWall time for provision, pull, init, fetch, load, compile, and warmup, recorded separatelyA single total hides which tier to fix
Cold-hit rateShare of requests that waited on a replica starting upThe user-facing cost of your scale-to-zero policy
TTFT p50 / p99 / p99.9Time to first token, including queue timeWhat users experience as "is it working?"
Inter-token latency p99Gap between consecutive streamed tokensCatches prefill interference and batch pressure
Preemptions / minSequences evicted or recomputed because KV memory ran outSilent latency and wasted compute
GoodputRequests per second that met their latency targetThroughput that doesn't count failures as work
Output driftRate at which identical greedy requests produce different tokensStability of the answer, not just the server

A playbook, one dimension at a time

Tier the bytes

Keep hot checkpoints in host RAM or local NVMe on GPU nodes, and use parallel ranged reads for everything else. Memory-map safetensors instead of unpickling. The goal is to make weight fetch a copy between local tiers, never a download from a region.

Snapshot the runtime, not only the weights

Persist compile caches and captured CUDA graphs alongside the checkpoint, keyed by model, GPU type, and parallelism layout. The dark bars in Fig. 1 are pure recomputation of something that hasn't changed since the last boot.

Warm pools sized by math

Endpoints with λ ≈ 1/T produce the most cold starts per hour. Keep a minimum replica there, let very quiet endpoints stay cold, and pre-warm on predictable signals like time of day or an upstream deploy.

Admit, don't just queue

Past about ρ = 0.8, an extra queued request mostly adds latency for everyone behind it. Shed or redirect load based on token budgets and deadlines, and give clients honest backpressure instead of silent timeouts that turn into retry storms.

Separate prefill from decode

Chunk long prefills so they can't stall every decoding stream in the batch, or split them onto their own pool. Measure inter-token latency separately from time to first token.

Budget KV cache like money

Price context length explicitly. Cap it per tier, share prefixes across requests, and alert on preemptions. At 128k context, a 70B node that fits hundreds of 2k chats fits about nine sequences.

Pin numerics where answers matter

Use batch-invariant kernels and a fixed parallelism layout for evaluations, audits, and anything cached or compared. Pay for determinism only on the paths that need it.

Hedge the tail

For fan-out and agent loops, send a backup request after the p95 delay and cancel the loser. A small increase in load buys a large cut in tail latency.

Where this leaves us

The web taught us to treat compute as stateless and disposable, and it worked because the state was small. A model server is the opposite: a few hundred gigabytes of state that takes minutes to assemble, shared by strangers whose presence changes your latency and even your answer. We keep applying stateless patterns to it and calling the results "AI being flaky."

My view is that runtime stability for models is a state-placement problem. It's about where the weights live, where the compiled runtime lives, and where each conversation's KV cache lives, and about routing work to the state instead of rebuilding the state for the work. The teams that get this right will ship systems that feel instant at 8:02 a.m. The others will keep buying faster GPUs and wondering why the spinner is still there.

I'm still working through how this changes once models are routinely swapped per request, which is the same statefulness question I keep coming back to in my notes on sidecar context architectures. If you run inference at scale and your numbers disagree with my model, I want to see them. That's the most useful email I can get.

References

  1. [1]Dean, J., & Barroso, L. A. (2013). The tail at scale. Communications of the ACM, 56(2), 74–80. link
  2. [2]Kwon, W., et al. (2023). Efficient memory management for large language model serving with PagedAttention. Proceedings of the 29th ACM Symposium on Operating Systems Principles (SOSP '23). link
  3. [3]Fu, Y., et al. (2024). ServerlessLLM: Low-latency serverless inference for large language models. 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI '24). link
  4. [4]Shahrad, M., et al. (2020). Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. USENIX Annual Technical Conference (ATC '20). link
  5. [5]Little, J. D. C. (1961). A proof for the queuing formula L = λW. Operations Research, 9(3), 383–387.
  6. [6]Kleinrock, L. (1975). Queueing Systems, Volume 1: Theory. Wiley.
  7. [7]NVIDIA. H100 Tensor Core GPU datasheet (SXM5: 80 GB HBM3, 3.35 TB/s). link
  8. [8]Grattafiori, A., et al. (2024). The Llama 3 herd of models. arXiv:2407.21783. link
  9. [9]Pope, R., et al. (2022). Efficiently scaling transformer inference. arXiv:2211.05102. link
  10. [10]Agrawal, A., et al. (2024). Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. OSDI '24. link
  11. [11]Zhong, Y., et al. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. OSDI '24. link
  12. [12]He, H., & Thinking Machines Lab (2025). Defeating nondeterminism in LLM inference. Connectionism. link
  13. [13]Hugging Face. Safetensors documentation. link