HBF for High-Throughput LLM Serving (UC Berkeley, FuriosaAI)
Original reporting by Semiconductor Engineering

High-bandwidth flash (HBF) refers to a type of non-volatile memory offering significantly greater data transfer rates than traditional flash storage, increasingly explored for its potential in large language model (LLM) serving. As LLMs grow in size and context lengths extend, the substantial memory demands for model weights and KV caches create critical bottlenecks, particularly exacerbated by agentic workloads requiring persistent context. HBF presents a compelling solution to expand accelerator memory capacity, yet its effective deployment is challenged by inherent access costs and concerns over limited write endurance.
Orchestrating HBF
A recent technical paper by researchers at UC Berkeley and FuriosaAI addresses these challenges head-on. They introduce a novel HBM-HBF-host hierarchical storage system combined with buffered cache-aware scheduling specifically designed for high-throughput agentic serving. Through trace-driven simulations, their work demonstrates remarkable improvements: HBF-augmented systems slashed LLM completion times by 36.1-87.0% compared to HBM-only setups, alongside significant energy savings of up to 55.8%. Crucially, their intelligent scheduling strategy extended HBF write lifetime from an estimated 4.77 years to 14.82 years, transforming a practical limitation into a viable long-term solution. This research underscores that strategic coordination of data placement and scheduling is paramount to unlocking HBF’s full potential for efficient and sustainable LLM serving.
The research from UC Berkeley and FuriosaAI presents a compelling case for High-Bandwidth Flash (HBF) as a critical component in the future of large language model (LLM) serving. By demonstrating how a carefully architected hierarchical storage system, coupled with intelligent, buffered cache-aware scheduling, can effectively mitigate HBF's inherent limitations – notably its access costs and write endurance – the paper offers a practical pathway to significantly boost LLM performance and energy efficiency. The observed reductions in completion time, reaching up to 87%, and the substantial extension of HBF lifespan to nearly 15 years are not merely incremental improvements; they represent a fundamental shift in how we can approach the scalability and cost-effectiveness of deploying advanced AI, especially for demanding agentic workloads. This work emphatically underscores the vital role of hardware-software co-design in unlocking the full potential of emerging memory technologies to overcome persistent bottlenecks.
Redefining LLM Infrastructure
The broader implications of this research extend far beyond mere technical optimization. By rendering HBF a viable, durable, and cost-effective option for memory expansion, these findings directly address the escalating hardware costs and energy demands associated with serving increasingly massive and complex LLMs. This breakthrough promises to democratize access to sophisticated AI capabilities, enabling more developers and organizations to deploy large models without prohibitive capital investments. Furthermore, by alleviating critical memory bottlenecks, it paves the way for the development of even larger models with longer context windows and more intricate agentic AI systems, pushing the boundaries of what AI can achieve in real-time, interactive environments. Ultimately, this work is poised to redefine the economic and architectural landscape of AI infrastructure, accelerating the pace of innovation and shaping the next generation of intelligent applications that were previously considered impractical.
Frequently asked questions
- What is High Bandwidth Flash (HBF) and how does it benefit large language models?
- High Bandwidth Flash (HBF) is a type of memory designed to expand the capacity available for storing large language model (LLM) weights and KV caches. As LLMs grow larger and handle longer contexts, traditional memory can become a bottleneck. HBF helps overcome this by providing additional high-capacity storage, reducing completion times and potentially improving energy efficiency for LLM serving by alleviating memory constraints.
- Why is memory capacity a significant challenge for efficient large language model serving?
- Large language models (LLMs) demand substantial memory to store their vast model weights and intermediate KV (Key-Value) caches. As these models increase in size and process longer conversational contexts, memory capacity and bandwidth become critical bottlenecks. This can severely limit serving performance and efficiency, especially for agentic workloads that involve repeated interactions and growing context sizes, necessitating solutions for memory expansion.
- What system design strategies enhance High Bandwidth Flash for LLM serving?
- Enhancing High Bandwidth Flash (HBF) for LLM serving involves coordinated system design and scheduling. Researchers propose a hierarchical storage system combining HBM, HBF, and host memory, alongside buffered cache-aware scheduling. These strategies are crucial for optimizing data placement and access patterns. This approach significantly reduces LLM completion times, improves energy efficiency, and extends HBF's write lifetime from years to over a decade, making its use practical.