Store Instead of Recompute: 10x Faster AI Inference, Beyond the Context Bottleneck
ET WQS High-Speed AI Semantic Storage
AI-Native KV Cache Storage · Distributed KV Cache Storage
WQS (WiDE Query Storage), designed from the ground up by ExponTech for large-scale AI inference, connects seamlessly to mainstream inference frameworks such as vLLM and SGLang through its native KV interface. It extends GPU HBM memory transparently with a distributed KV Cache storage pool, delivering unlimited capacity scaling and cluster-wide global sharing — a storage foundation built for long-context, high-concurrency inference. Store instead of recompute — break through the context bottleneck.
20x
TTFT speedup (100K-token input)
22x
Token throughput gain
<200μs
End-to-end read latency
97–98%
Cache hit rate in production
100%
DPU-native offload (optional)
Three Innovation Pillars
Fully native design built around the I/O characteristics of KV Cache — not a patch on legacy storage
AI-Native KV Interface & KV Engine
Native "KV → network → KV" path: 2 hops, 0 kernel crossings, fully user-space RDMA zero-copy
Manages raw NVMe SSDs directly with unified pooling, bypassing the lengthy file-system I/O path and kernel overhead
Flat full-bandwidth across 16K–10M small I/O, while DFS bandwidth collapses ~10× on small I/O
CacheBlend non-prefix reuse: matches cached KV at any position in a sequence, effective hit rate up to 98%+
Layer-Wise Pipelined Loading
Prefetches layer i+1 KV while the GPU computes layer i — I/O latency fully hidden behind compute
Per-layer KV load time down to <10μs; the SSD-vs-HBM latency gap no longer bottlenecks end-to-end inference
Block granularity down to 16 tokens (16× finer than conventional), yielding finer matching and higher hit rates
DPU-Native Architecture (Optional)
Software runs 100% on NVIDIA BlueField DPU + JBOF — zero host CPU involvement in the storage path
No dedicated storage servers needed; 1 WQS node ≈ 3 traditional storage servers of equivalent performance
Perfectly aligned with the NVIDIA CMX tiered caching architecture, serving as the G3.5 shared KV pool layer
System Architecture
End-to-end layering from application to media: zero-intrusion access for inference frameworks, end-to-end latency <200μs
System Architecture
Measured Performance
Joint testing on production-equivalent hardware: lab stress scenario and real-workload A/B comparison
Lab Stress Scenario: Full Recompute vs. Full Hit
DeepSeek-R1-0528 · 8× NVIDIA H20 · vLLM + LMCache · 100K-token input
MetricGPU recomputeWQS full hitGain
TTFT (time to first token)323 s16s20x
Token throughput4,194 tok/s90,000 tok/s22x
Inter-token latency (ITL)baseline≈ 1/1212x
End-to-end read latency—<200μs—
Real-Workload A/B Comparison: Multi-turn Agents at High Concurrency
GLM 4.7 355B · 30K-token system prompt · multi-turn sessions · 30–100 concurrent users · HBM-only vs. HBM + WQS
ConcurrencyTTFT · HBM onlyTTFT · HBM+WQSThroughputHit rate
3022.38s5.03s4.5x97%
60179.56s42.19s4.4x97%
100399.35s88.03s4.7x97%
HBM-only hits an "eviction cliff" at 30 concurrent users — TTFT degrades catastrophically to 399 s. WQS absorbs the overflowing KV Cache beyond HBM, keeping performance flat and predictable: a stable 4.4–4.7× throughput gain and 76–78% TTFT reduction on real workloads.
Integration with vLLM + LMCache
Zero intrusion into the inference stack: keep your existing vLLM + LMCache deployment and attach the WQS pool in three steps
Step 1 · Start the WQS storage cluster
Deploy the wqs_vault / wqsd storage services (RDMA transport, NVMe pool capacity, CPU pinning) to serve KV Cache requests.
Step 2 · Attach the WQS L2 adapter to LMCache Server
Launch lmcache server with --l2-adapter type wqs_store for RDMA-direct access to the pool, with a hugepage memfd-backed L1.
Step 3 · Hook vLLM into LMCache
vLLM connects through LMCacheConnectorV1 with zero code changes; prefix KV automatically tiers into WQS, shared cluster-wide.
Integration with vLLM + LMCache
Flexible Deployment Options
One software kernel, multiple form factors — compatible with vLLM / SGLang / LMCache / Mooncake / Dynamo / NVIDIA CMX
Software-Only
Standard x86 servers with local NVMe SSDs and 200G+ RoCE NICs. Minimum footprint: a single node with 16 cores and 48 GB RAM; no dedicated metadata disk required. Lowest barrier to deploy.
DPU + JBOF (High Performance)
WQS runs entirely on NVIDIA BlueField DPUs, freeing 100% of host CPUs for inference. A 2U chassis with 4× DPUs and 26× NVMe delivers up to 800Gbps × 4 bandwidth at a much lower $/TB.
Turnkey Appliance
Pre-installed WQS integrated storage servers — out-of-the-box deployment, deeply tuned and certified hardware, unified operations console, with SSD supply secured via strategic vendor partnerships.
Store Instead of Recompute: 10x Faster AI Inference Beyond the Context Bottleneck - ExponTech