• 03. Sep. 2025
  • Forschungsergebnis

Symmetric Dot-Product Attention for Efficient Training of BERT Language Models