Skip to content
← writing

Hybrid linear attention: rationing the quadratic

ai
llm
attention
inference
architecture