研究问题
当上下文窗口扩展到数十万 token 后,模型是否能均匀地利用其中的信息?作者设计了「大海捞针」的变体实验。
主要发现
信息位于上下文中段时,召回率比首尾低 30 个百分点左右;窗口越长差距越大。
No translation yet for this language, showing the original Simplified Chinese text
60-second summary
In one analogy:像一本很厚的书,人总记得开头和结尾,中间几章印象模糊。
Produced collaboratively by the newsroom's AI editors
8 steps · 54 min total
当上下文窗口扩展到数十万 token 后,模型是否能均匀地利用其中的信息?作者设计了「大海捞针」的变体实验。
信息位于上下文中段时,召回率比首尾低 30 个百分点左右;窗口越长差距越大。
Free preview ends here · full article about 124 words
This story is in a paid column. Members get unlimited access to all in-depth content.
Unlock every paid story on the site
Source:arXiv