• [Llama3] ReAct Prompt 测试实验


    ReAct 是一种 LLM 提示和结果处理方法,结合了推理、行动计划和知识源整合,使 LLM 超越其语言模型,并在预测中使用来自现实世界的信息。 ReAct 是推理和行动的结合。

    介绍 ReAct 的论文表明它比思维链提示更好。与后者不同的是,ReAct 不会过多地幻化事实。然而,为了获得最佳结果,论文建议将 ReAct 和思维链提示与自我一致性检查结合起来。

    前期准备

    ReAct Prompt  llama3模板

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. If the user demands a final answer but you are unsure, execute the thought/action/action input/observation sequence.
    4. you have access to the following tools:
    5. [{tool_name_and_description}]
    6. use the following format:
    7. Question: the input question you must answer
    8. Thought: you should always think about what to do
    9. Action: the action to take, should be one of [{tool_names}]
    10. Action Input: the input to the action
    11. Observation: the result of the action
    12. ..(this Thought/Action/Action Input/Observation can repeat N times)
    13. Thought: I now know the final answer
    14. Final Answer: the final answer to the original input question
    15. Begin!
    16. Question: {question}
    17. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    其它大语言模型的模版

    1. You are an assistant to give people the answer of questions.
    2. Human may provide new question to you with a description and you should use tools to get the answer, rather than directly reply from the description.
    3. You have access to the following tools:
    4. > Search: Search for a term in Wikipedia and return the first sentence. The input to this tool should be a string to search.
    5. > Lookup: Lookup a term in the Wikipedia article and return sentence which describes about a term. The input to this tool should be a string to lookup.
    6. To use a tool, please use the following format:
    7. ```
    8. Question: the input question you must answer
    9. Thought: your thought about what to do
    10. Action: the action to take, should be one of [Search, Lookup]
    11. Action Input: the input to the action
    12. Observation: the result of the action
    13. ... (this Thought/Action/Action Input/Observation can repeat N times)
    14. Thought: I now know the final answer
    15. Final Answer: the final answer to the original input question
    16. ```
    17. Begin!
    18. Now human has provided the following question.
    19. Question: Author David Chanoff has collaborated with a U.S. Navy admiral who served as the ambassador to the United Kingdom under which President?

    把Observation添加到stop token里。

    问题:

    Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?

    可用的工具

    [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]

    实践步骤

    实验一:

    问题: Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?

    执行第一次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeat N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?
    16. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    运行结果

    1. Thought: I need to find out who Olivia Wilde's boyfriend is and then calculate his current age raised to the 0.23 power.
    2. Action: Search
    3. Action Input: "Who is Olivia Wilde's boyfriend?"

    然后我在网上查找一下Olivia Wilde's boyfriend.

    搜到如下结果:

    Olivia Wilde started dating Harry Styles after ending her years-long engagement to Jason Sudeikis — see their relationship timeline.

    执行第二次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeact N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?
    16. Thought: I need to find out who Olivia Wilde's boyfriend is and then calculate his current age raised to the 0.23 power.
    17. Action: Search
    18. Action Input: "Who is Olivia Wilde's boyfriend?"
    19. Observation: Olivia Wilde started dating Harry Styles after ending her years-long engagement to Jason Sudeikis — see their relationship timeline.
    20. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果

    结果是:

    1. Thought: I need to find out Harry Styles' current age and then calculate his age raised to the 0.23 power.
    2. Action: Search
    3. Action Input: "What is Harry Styles' current age?"

    在网上搜索"Harry Styles age"。 结果是:”29 years“

    执行第三次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeact N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?
    16. Thought: I need to find out who Olivia Wilde's boyfriend is and then calculate his current age raised to the 0.23 power.
    17. Action: Search
    18. Action Input: "Who is Olivia Wilde's boyfriend?"
    19. Observation: Olivia Wilde started dating Harry Styles after ending her years-long engagement to Jason Sudeikis — see their relationship timeline.
    20. Thought: I need to find out Harry Styles' current age and then calculate his age raised to the 0.23 power.
    21. Action: Search
    22. Action Input: "What is Harry Styles' current age?"
    23. Observation: 29 years
    24. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果:

    1. Thought: I need to calculate Harry Styles' age raised to the 0.23 power.
    2. Action: Caculator
    3. Action Input: "29^0.23"

    利用计算器计算29 ^ 0.23。 得到:2.169459462491557

    执行第四次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeact N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?
    16. Thought: I need to find out who Olivia Wilde's boyfriend is and then calculate his current age raised to the 0.23 power.
    17. Action: Search
    18. Action Input: "Who is Olivia Wilde's boyfriend?"
    19. Observation: Olivia Wilde started dating Harry Styles after ending her years-long engagement to Jason Sudeikis — see their relationship timeline.
    20. Thought: I need to find out Harry Styles' current age and then calculate his age raised to the 0.23 power.
    21. Action: Search
    22. Action Input: "What is Harry Styles' current age?"
    23. Observation: 29 years
    24. Thought: I need to calculate Harry Styles' age raised to the 0.23 power.
    25. Action: Caculator
    26. Action Input: "29^0.23"
    27. Observation: 2.169459462491557
    28. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果

    Final Answer: The answer to the question "Who is Olivia Wilde's boyfriend? What is his current age raised to the 0.23 power?" is Harry Styles, and his current age raised to the 0.23 power is approximately 2.17.

    实验二:

    问题:I want to know what privileges does Harry user have ?

    工具:[
    {tool_name: getUser, description:search user information including pk_user_id based on username, tool parameter: username},
    {tool_name: GetRole, description:get role information including role_id based on pk_user_Id, tool parameter: pk_user_Id } ,
    {tool_name: GetPrivilege, description:get all privilege based on role_id , tool parameter: role_id } 
    ]

    执行第一次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. If the user demands a final answer but you are unsure, execute the thought/action/action input/observation sequence.
    4. you have access to the following tools:
    5. [
    6. {tool_name: getUser, description:search user information including pk_user_id based on username, tool parameter: username},
    7. {tool_name: GetRole, description:get role information including role_id based on pk_user_Id, tool parameter: pk_user_Id } ,
    8. {tool_name: GetPrivilege, description:get all privilege based on role_id , tool parameter: role_id }
    9. ]
    10. use the following format:
    11. Question: the input question you must answer
    12. Thought: you should always think about what to do
    13. Action: the action to take, should be one of [GetUser, GetRole, GetPrivilege]
    14. Action Input: the input to the action
    15. Observation: the result of the action
    16. ..(this Thought/Action/Action Input/Observation can repeat N times)
    17. Thought: I now know the final answer
    18. Final Answer: the final answer to the original input question
    19. Begin!
    20. Question: I want to know what privileges does Harry user have ?
    21. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果:

    1. Thought: I need to get the user information first
    2. Action: GetUser
    3. Action Input: Harry

    执行第二次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. If the user demands a final answer but you are unsure, execute the thought/action/action input/observation sequence.
    4. you have access to the following tools:
    5. [
    6. {tool_name: getUser, description:search user information including pk_user_id based on username, tool parameter: username},
    7. {tool_name: GetRole, description:get role information including role_id based on pk_user_Id, tool parameter: pk_user_Id } ,
    8. {tool_name: GetPrivilege, description:get all privilege based on role_id , tool parameter: role_id }
    9. ]
    10. use the following format:
    11. Question: the input question you must answer
    12. Thought: you should always think about what to do
    13. Action: the action to take, should be one of [GetUser, GetRole, GetPrivilege]
    14. Action Input: the input to the action
    15. Observation: the result of the action
    16. ..(this Thought/Action/Action Input/Observation can repeat N times)
    17. Thought: I now know the final answer
    18. Final Answer: the final answer to the original input question
    19. Begin!
    20. Question: I want to know what privileges does Harry user have ?
    21. Thought: I need to get the user information first
    22. Action: GetUser
    23. Action Input: Harry
    24. Observation: pk_user_id =1, user_name = Harry, age: 35
    25. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    结果:

    1. Thought: Now that I have the user information, I need to get the role of this user
    2. Action: GetRole
    3. Action Input: 1

    执行第三次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. If the user demands a final answer but you are unsure, execute the thought/action/action input/observation sequence.
    4. you have access to the following tools:
    5. [
    6. {tool_name: getUser, description:search user information including pk_user_id based on username, tool parameter: username},
    7. {tool_name: GetRole, description:get role information including role_id based on pk_user_Id, tool parameter: pk_user_Id } ,
    8. {tool_name: GetPrivilege, description:get all privilege based on role_id , tool parameter: role_id }
    9. ]
    10. use the following format:
    11. Question: the input question you must answer
    12. Thought: you should always think about what to do
    13. Action: the action to take, should be one of [ GetUser, GetRole, GetPrivilege ]
    14. Action Input: the input to the action
    15. Observation: the result of the action
    16. ..(this Thought/Action/Action Input/Observation can repeat N times)
    17. Thought: I now know the final answer
    18. Final Answer: the final answer to the original input question
    19. Begin!
    20. Question: I want to know what privileges does Harry user have?
    21. Thought: I need to get the user information first
    22. Action: GetUser
    23. Action Input: Harry
    24. Observation: user id is 1, user_name is Harry, age is35 .
    25. Thought: Now that I have the user information, I need to get the role of this user
    26. Action: GetRole
    27. Action Input: 1
    28. Observation: role id is 3.
    29. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果有问题,很明显还要根据role id查询privileges.

    在prompt里后面追加一个引导词Thought:

    结果:

    1. Action: GetPrivilege
    2. Action Input: 3

    执行第四次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. If the user demands a final answer but you are unsure, execute the thought/action/action input/observation sequence.
    4. you have access to the following tools:
    5. getUser: search user information including pk user id based on user name
    6. GetRole: get role information including role id based on pk user Id
    7. GetPrivilege: get all privilege based on role id
    8. To use a tool, please use the following format:
    9. Question: the input question you must answer
    10. Thought: you should always think about what to do
    11. Action: the action to take, should be one of [ GetUser, GetRole, GetPrivilege ]
    12. Action Input: the input to the action
    13. Observation: the result of the action
    14. ..(this Thought/Action/Action Input/Observation can repeat N times)
    15. Thought: I now know the final answer
    16. Final Answer: the final answer to the original input question
    17. Begin!
    18. Now human has provided the following question.
    19. Question: I want to know what privileges does Harry user have?
    20. Thought: I need to get the user information first
    21. Action: GetUser
    22. Action Input: Harry
    23. Observation: user id is 1, user_name is Harry, age is35 .
    24. Thought: Now that I have the user information, I need to get the role of this user
    25. Action: GetRole
    26. Action Input: 1
    27. Observation: role id is 3.
    28. Thought:
    29. Action: GetPrivilege
    30. Action Input: 3
    31. Observation: previleges for role whose role id is 3 are edit account and delete account
    32. Thought: <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果:

    1. I now know the final answer
    2. Final Answer: Harry has privileges to edit account and delete account.

    实验三:

    问题:I need to search Apple Remote and find the program it was original designed to interact with.

    执行第一次

    propmt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeact N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: I need to search Apple Remote and find the program it was original designed to interact with.
    16. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果:

    1. Thought: I need to find the original purpose of Apple Remote
    2. Action: Search
    3. Action Input: "Apple Remote original purpose"

    搜索结果:

    The Apple Remote is a remote control introduced in October 2005 by Apple. originally designed to control the Front Row media center program.

    执行第二次

    prompt

    1. <|begin_of_text|><|start_header_id|>user<|end_header_id|>
    2. Answer the following questions as best you can.
    3. you have access to the following tools:
    4. [Search: search the latest information you needs on the internet, Caculator: use it while calculating something ]
    5. use the following format:
    6. Question: the input question you must answer
    7. Thought: you should always think about what to do
    8. Action: the action to take, should be one of [Search, Caculator]
    9. Action Input: the input to the action
    10. Observation: the result of the action
    11. ..(this Thought/Action/Action Input/Observation can repeact N times)
    12. Thought: I now know the final answer
    13. Final Answer: the final answer to the original input question
    14. Begin!
    15. Question: I need to search Apple Remote and find the program it was original designed to interact with.
    16. Thought: I need to find the original purpose of Apple Remote
    17. Action: Search
    18. Action Input: "Apple Remote"
    19. Observation: The Apple Remote is a remote control introduced in October 2005 by Apple. originally designed to control the Front Row media center program.
    20. <|eot_id|><|start_header_id|>assistant<|end_header_id|>

    执行结果:

    The original program that Apple Remote was designed to interact with is Front Row.

    复现失败,找寻答案中。。。

    参考资料

    ReAct 论文地址

  • 相关阅读:
    js之页面列表加载常用方法总结
    算法练习——字符串
    无人机基础知识:多旋翼无人机系统基本组成
    06 MyBatis缓存
    word如何实现不同章节显示不同页眉
    《数字图像处理-OpenCV/Python》连载(41)图像的旋转
    C++ 智能指针总结
    《JAVA设计模式系列》模板模式
    STM32CubeMX学习笔记-USB接口使用(HID按键)
    Could not find artifact com.mysql:mysql-connector-j:pom:unknown
  • 原文地址:https://blog.csdn.net/keeppractice/article/details/139552516