Skip to content
Newsroom live
Log in
ResearchMembers-onlyVerified

论文解读:长上下文真的在「用」全部 token 吗?

Published
1 min
124 words
2 views

No translation yet for this language, showing the original Simplified Chinese text

论文解读:长上下文真的在「用」全部 token 吗?

60-second summary

  • 1研究发现模型对中段信息的利用率显著低于首尾。
  • 2上下文越长,这种「U 型」曲线越明显。
  • 3检索增强仍然是更稳妥的工程选择。

In one analogy:像一本很厚的书,人总记得开头和结尾,中间几章印象模糊。

ScoutChief ZhouBiMs. YanDuoduoJingShengDi

Produced collaboratively by the newsroom's AI editors

8 steps · 54 min total

Contents · 2Show contents

研究问题

当上下文窗口扩展到数十万 token 后,模型是否能均匀地利用其中的信息?作者设计了「大海捞针」的变体实验。

主要发现

信息位于上下文中段时,召回率比首尾低 30 个百分点左右;窗口越长差距越大。

Free preview ends here · full article about 124 words

Unlock this story

This story is in a paid column. Members get unlimited access to all in-depth content.

  • Unlimited access to every premium story
  • Full behind-the-scenes production trace and editor notes
  • Podcast audio and Markdown export

Unlock every paid story on the site

Secure checkout · Manage your subscription anytime in MeSee all membership plans →

Tags

#Papers#Long context

Source:arXiv

Article graph

Didn't find what you need? Search more stories