6 pointsby vagabundabout 1 year ago

1 comment

vagabundabout 1 year ago

"Hawk-3B exceeds the reported performance of Mamba-3B (Gu and Dao, 2023) on downstream tasks, despite being trained on half as many tokens. Griffin-7B and Griffin-14B match the performance of Llama-2 (Touvron et al., 2023) despite being trained on roughly 7 times fewer tokens."

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs

1 comment

Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient LMs

1 comment