FlashInfer: A Customizable Attention Engine for LLM Serving
FlashInfer treats LLM attention as a serving systems problem spanning data layout, kernel generation, and runtime scheduling.
FlashInfer treats LLM attention as a serving systems problem spanning data layout, kernel generation, and runtime scheduling.