ReAct架构通过推理与行动循环将大模型转化为具备纠错能力的Agent
大模型由于是预测下一个token的概率模型,无法感知真实世界且难以发现错误,容易产生幻觉。ReAct(Reason + Act)架构通过构建“想一下、动一下、看一眼”的闭环,将文本生成器改造为能与环境互动的执行体。
相比于仅在内部推演、一旦第一步出错便会全盘跑偏的Chain-of-Thought(CoT),ReAct让推理和行动交替进行:模型先通过Thought判断局面,执行Action调用外部工具,再通过Observation获取工具返回结果。
这种架构将外部反馈作为纠偏锚点。例如在查询公司财报时,若模型在Thought阶段误判CEO人选,但在Action阶段通过搜索API获得的Observation与之矛盾,模型会在下一个循环中自动修正方向,从而提升复杂任务的完成率。
开发者无需从零训练模型,仅需在Prompt层搭建强制格式框架即可实现ReAct。典型的循环Prompt如下:
Answer the following questions as best you can. You have access to the following tools:
[Search: a tool to search the web]
Use the following format:
Question: the input question you must answer
Thought: you should always think about what to do
Action: the action to take, should be one of [Search]
Action Input: the input to the action
Observation: the result of the action
... (this Thought/Action/Observation can repeat N times)
Thought: I now know the final answer
Final Answer: the final answer to the original input question
ReAct的天花板取决于Observation阶段提供的信息质量而非模型推理能力,如果工具返回垃圾数据,模型仍会被误导。这意味着开发重心正从打磨Prompt转向搭建精准的工具集。
ReAct是Agent走向工程化的分水岭,在不增加参数量的前提下提升了模型应对动态、未知信息的稳定性。未来这类循环将更加轻量化与自动化,从依赖长篇Prompt引导转向融入模型自身的底层能力。
