optimizing-attention-flash：优化Transformer注意力，提速2-4倍并大幅节省内存。

hub · 2026-02-19 17:40:11 · 22 次点击 · 0 条评论

Optimizes transformer attention with Flash Attention for 2-4x speedup and 10-20x memory reduction. Use when training/running transformers with long sequences (>512 tokens), encountering GPU memory issues with attention, or need faster inference. Supports PyTorch native SDPA, flash-attn library, H100 FP8, and sliding window attention.

技能包地址：https://skillsmp.com/skills/davila7-claude-code-templates-cli-tool-components-skills-ai-research-optimization-flash-attention-skill-md

22 次点击 ∙ 0 人收藏

登录后收藏

0 条回复