Your browser does not fully support modern features. Please upgrade for a smoother experience.
Submitted Successfully!
Thank you for your contribution! You can also upload a video entry or images related to this topic. For video creation, please contact our Academic Video Service.
Version Summary Created by Modification Content Size Created at Operation
1 handwiki Dean Liu -- 3971 2022-11-22 01:35:17

Video Upload Options

We provide professional Academic Video Service to translate complex research into visually appealing presentations. Would you like to try it?
Cite
If you have any further questions, please contact Encyclopedia Editorial Office.
HandWiki. AI Control Problem. Encyclopedia. Available online: https://encyclopedia.pub/entry/35791 (accessed on 28 September 2026).
HandWiki. AI Control Problem. Encyclopedia. Available at: https://encyclopedia.pub/entry/35791. Accessed September 28, 2026.
HandWiki. "AI Control Problem" Encyclopedia, https://encyclopedia.pub/entry/35791 (accessed September 28, 2026).
HandWiki. (2022, November 22). AI Control Problem. In Encyclopedia. https://encyclopedia.pub/entry/35791
HandWiki. "AI Control Problem." Encyclopedia. Web. 22 November, 2022.
AI Control Problem
Edit

In artificial intelligence (AI) and philosophy, the AI control problem is the issue of how to build AI systems such that they will aid their creators, and avoid inadvertently building systems that will harm their creators. One particular concern is that humanity will have to solve the control problem before a superintelligent AI system is created, as a poorly designed superintelligence might rationally decide to seize control over its environment and refuse to permit its creators to modify it after launch. In addition, some scholars argue that solutions to the control problem, alongside other advances in AI safety engineering, might also find applications in existing non-superintelligent AI. Major approaches to the control problem include alignment, which aims to align AI goal systems with human values, and capability control, which aims to reduce an AI system's capacity to harm humans or gain control. Capability control proposals are generally not considered reliable or sufficient to solve the control problem, but rather as potentially valuable supplements to alignment efforts.

artificial intelligence control problem philosophy

References

  1. Bostrom, Nick (2014). Superintelligence: Paths, Dangers, Strategies (First ed.). ISBN 978-0199678112. 
  2. "Stephen Hawking: 'Transcendence looks at the implications of artificial intelligence – but are we taking AI seriously enough?'". The Independent (UK). https://www.independent.co.uk/news/science/stephen-hawking-transcendence-looks-at-the-implications-of-artificial-intelligence--but-are-we-taking-ai-seriously-enough-9313474.html. 
  3. "Stephen Hawking warns artificial intelligence could end mankind". BBC. 2 December 2014. https://www.bbc.com/news/technology-30290540. 
  4. "Anticipating artificial intelligence". Nature 532 (7600): 413. 26 April 2016. doi:10.1038/532413a. PMID 27121801. Bibcode: 2016Natur.532Q.413..  https://dx.doi.org/10.1038%2F532413a
  5. Russell, Stuart; Norvig, Peter (2009). "26.3: The Ethics and Risks of Developing Artificial Intelligence". Artificial Intelligence: A Modern Approach. Prentice Hall. ISBN 978-0-13-604259-4. 
  6. Dietterich, Thomas; Horvitz, Eric (2015). "Rise of Concerns about AI: Reflections and Directions". Communications of the ACM 58 (10): 38–40. doi:10.1145/2770869. http://research.microsoft.com/en-us/um/people/horvitz/CACM_Oct_2015-VP.pdf. Retrieved 14 June 2016. 
  7. Russell, Stuart (2014). "Of Myths and Moonshine". http://edge.org/conversation/the-myth-of-ai#26015. 
  8. "Google developing kill switch for AI". BBC News. 8 June 2016. https://www.bbc.com/news/technology-36472140. 
  9. "'Press the big red button': Computer experts want kill switch to stop robots from going rogue". Washington Post. https://www.washingtonpost.com/news/morning-mix/wp/2016/06/09/press-the-big-red-button-computer-experts-want-kill-switch-to-stop-robots-from-going-rogue/. 
  10. "DeepMind Has Simple Tests That Might Prevent Elon Musk's AI Apocalypse". Bloomberg.com. 11 December 2017. https://www.bloomberg.com/news/articles/2017-12-11/deepmind-has-simple-tests-that-might-prevent-elon-musk-s-ai-apocalypse. 
  11. "Alphabet's DeepMind Is Using Games to Discover If Artificial Intelligence Can Break Free and Kill Us All" (in en). Fortune. http://fortune.com/2017/12/12/alphabet-deepmind-ai-safety-musk-games/. 
  12. "Specifying AI safety problems in simple environments | DeepMind". https://deepmind.com/blog/specifying-ai-safety-problems/. 
  13. Gabriel, Iason (1 September 2020). "Artificial Intelligence, Values, and Alignment" (in en). Minds and Machines 30 (3): 411–437. doi:10.1007/s11023-020-09539-2. ISSN 1572-8641. https://link.springer.com/article/10.1007/s11023-020-09539-2. Retrieved 7 February 2021. 
  14. Russell, Stuart (October 8, 2019). Human Compatible: Artificial Intelligence and the Problem of Control. United States: Viking. ISBN 978-0-525-55861-3. OCLC 1083694322. https://archive.org/details/humancompatiblea0000russ. 
  15. Yudkowsky, Eliezer (2011). "Complex Value Systems in Friendly AI". Artificial General Intelligence. Lecture Notes in Computer Science. 6830. pp. 388–393. doi:10.1007/978-3-642-22887-2_48. ISBN 978-3-642-22886-5.  https://dx.doi.org/10.1007%2F978-3-642-22887-2_48
  16. Leike, Jan; Krueger, David; Everitt, Tom; Martic, Miljan; Maini, Vishal; Legg, Shane (19 November 2018). "Scalable agent alignment via reward modeling: a research direction". arXiv:1811.07871 [cs.LG]. //arxiv.org/archive/cs.LG
  17. Ortega, Pedro; Maini, Vishal; DeepMind Safety Team (27 September 2018). "Building safe artificial intelligence: specification, robustness, and assurance" (in en). https://medium.com/@deepmindsafetyresearch/building-safe-artificial-intelligence-52f5f75058f1. 
  18. Hubinger, Evan; van Merwijk, Chris; Mikulik, Vladimir; Skalse, Joar; Garrabrant, Scott (11 June 2019). "Risks from Learned Optimization in Advanced Machine Learning Systems". arXiv:1906.01820 [cs.AI]. //arxiv.org/archive/cs.AI
  19. Ecoffet, Adrien; Clune, Jeff; Lehman, Joel (1 July 2020). "Open Questions in Creating Safe Open-ended AI: Tensions Between Control and Creativity". Artificial Life Conference Proceedings 32: 27–35. doi:10.1162/isal_a_00323. https://www.mitpressjournals.org/doi/abs/10.1162/isal_a_00323. 
  20. Christian, Brian (2020) (in en). The Alignment Problem: Machine Learning and Human Values. W.W. Norton. ISBN 978-0-393-63582-9. https://www.google.co.uk/books/edition/The_Alignment_Problem/VmJIzQEACAAJ?hl=en. Retrieved 2021-02-07. 
  21. Krakovna, Victoria; Legg, Shane. "Specification gaming: the flip side of AI ingenuity". https://deepmind.com/blog/article/Specification-gaming-the-flip-side-of-AI-ingenuity. 
  22. Clark, Jack; Amodei, Dario (22 December 2016). "Faulty Reward Functions in the Wild" (in en). https://openai.com/blog/faulty-reward-functions/. 
  23. Christiano, Paul (11 September 2019). "Conversation with Paul Christiano". AI Impacts. https://aiimpacts.org/conversation-with-paul-christiano/. 
  24. Serban, Alex; Poll, Erik; Visser, Joost (12 June 2020). "Adversarial Examples on Object Recognition: A Comprehensive Survey". ACM Computing Surveys 53 (3): 66:1–66:38. doi:10.1145/3398394. ISSN 0360-0300. https://dl.acm.org/doi/abs/10.1145/3398394. Retrieved 7 February 2021. 
  25. Kohli, Pushmeet; Dvijohtham, Krishnamurthy; Uesato, Jonathan; Gowal, Sven. "Towards Robust and Verified AI: Specification Testing, Robust Training, and Formal Verification". https://deepmind.com/blog/article/robust-and-verified-ai. 
  26. Christiano, Paul; Leike, Jan; Brown, Tom; Martic, Miljan; Legg, Shane; Amodei, Dario (13 July 2017). "Deep Reinforcement Learning from Human Preferences". arXiv:1706.03741 [stat.ML]. //arxiv.org/archive/stat.ML
  27. Amodei, Dario; Olah, Chris; Steinhardt, Jacob; Christiano, Paul; Schulman, John; Mané, Dan (25 July 2016). "Concrete Problems in AI Safety". arXiv:1606.06565 [cs.AI]. //arxiv.org/archive/cs.AI
  28. Amodei, Dario; Christiano, Paul; Ray, Alex (13 June 2017). "Learning from Human Preferences" (in en). https://openai.com/blog/deep-reinforcement-learning-from-human-preferences/. 
  29. Christiano, Paul; Shlegeris, Buck; Amodei, Dario (19 October 2018). "Supervising strong learners by amplifying weak experts". arXiv:1810.08575 [cs.LG]. //arxiv.org/archive/cs.LG
  30. Irving, Geoffrey; Christiano, Paul; Amodei, Dario; OpenAI (October 22, 2018). "AI safety via debate". arXiv:1805.00899 [stat.ML]. //arxiv.org/archive/stat.ML
  31. Banzhaf, Wolfgang; Goodman, Erik; Sheneman, Leigh; Trujillo, Leonardo; Worzel, Bill (May 2020) (in en). Genetic Programming Theory and Practice XVII. Springer Nature. ISBN 978-3-030-39958-0. https://www.google.co.uk/books/edition/Genetic_Programming_Theory_and_Practice/2U7iDwAAQBAJ?hl=en&gbpv=1&pg=PA181&printsec=frontcover#v=onepage&q&f=false. Retrieved 2021-02-07. 
  32. Stiennon, Nisan; Ziegler, Daniel; Lowe, Ryan; Wu, Jeffrey; Voss, Chelsea; Christiano, Paul; Ouyang, Long (September 4, 2020). "Learning to Summarize with Human Feedback". https://openai.com/blog/learning-to-summarize-with-human-feedback/. 
  33. Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (12 November 2016). "Cooperative Inverse Reinforcement Learning". Neural Information Processing Systems. 
  34. Everitt, Tom; Lea, Gary; Hutter, Marcus (21 May 2018). "AGI Safety Literature Review". 1805.01109. 
  35. Demski, Abram; Garrabrant, Scott (6 October 2020). "Embedded Agency". arXiv:1902.09469 [cs.AI]. //arxiv.org/archive/cs.AI
  36. Everitt, Tom; Ortega, Pedro A.; Barnes, Elizabeth; Legg, Shane (6 September 2019). "Understanding Agent Incentives using Causal Influence Diagrams. Part I: Single Action Settings". arXiv:1902.09980 [cs.AI]. //arxiv.org/archive/cs.AI
  37. Everitt, Tom; Hutter, Marcus (20 August 2019). "Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective". arXiv:1908.04734 [cs.AI]. //arxiv.org/archive/cs.AI
  38. Everitt, Tom; Filan, Daniel; Daswani, Mayank; Hutter, Marcus (10 May 2016). "Self-Modification of Policy and Utility Function in Rational Agents". arXiv:1605.03142 [cs.AI]. //arxiv.org/archive/cs.AI
  39. Leike, Jan; Taylor, Jessica; Fallenstein, Benya (25 June 2016). "A formal solution to the grain of truth problem". Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence. UAI'16 (AUAI Press): 427–436. ISBN 9780996643115. https://dl.acm.org/doi/10.5555/3020948.3020993. Retrieved 7 February 2021. 
  40. Garrabrant, Scott; Benson-Tilsen, Tsvi; Critch, Andrew; Soares, Nate; Taylor, Jessica (7 December 2020). "Logical Induction". arXiv:1609.03543 [cs.AI]. //arxiv.org/archive/cs.AI
  41. Montavon, Grégoire; Samek, Wojciech; Müller, Klaus Robert (2018). "Methods for interpreting and understanding deep neural networks" (in English). Digital Signal Processing: A Review Journal 73: 1–15. doi:10.1016/j.dsp.2017.10.011. ISSN 1051-2004.  https://dx.doi.org/10.1016%2Fj.dsp.2017.10.011
  42. Yampolskiy, Roman V. "Unexplainability and Incomprehensibility of AI." Journal of Artificial Intelligence and Consciousness 7.02 (2020): 277-291.
  43. Hadfield-Menell, Dylan; Dragan, Anca; Abbeel, Pieter; Russell, Stuart (15 June 2017). "The Off-Switch Game". arXiv:1611.08219 [cs.AI]. //arxiv.org/archive/cs.AI
  44. Orseau, Laurent; Armstrong, Stuart (25 June 2016). "Safely interruptible agents". Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence. UAI'16 (AUAI Press): 557–566. ISBN 9780996643115. https://dl.acm.org/doi/10.5555/3020948.3021006. Retrieved 7 February 2021. 
  45. Soares, Nate, et al. "Corrigibility." Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence. 2015.
  46. Chalmers, David (2010). "The singularity: A philosophical analysis". Journal of Consciousness Studies 17 (9–10): 7–65. 
  47. Bostrom, Nick (2014). "Chapter 10: Oracles, genies, sovereigns, tools (page 145)". Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press. ISBN 9780199678112. "An oracle is a question-answering system. It might accept questions in a natural language and present its answers as text. An oracle that accepts only yes/no questions could output its best guess with a single bit, or perhaps with a few extra bits to represent its degree of confidence. An oracle that accepts open-ended questions would need some metric with which to rank possible truthful answers in terms of their informativeness or appropriateness. In either case, building an oracle that has a fully domain-general ability to answer natural language questions is an AI-complete problem. If one could do that, one could probably also build an AI that has a decent ability to understand human intentions as well as human words." 
  48. Armstrong, Stuart; Sandberg, Anders; Bostrom, Nick (2012). "Thinking Inside the Box: Controlling and Using an Oracle AI". Minds and Machines 22 (4): 299–324. doi:10.1007/s11023-012-9282-2.  https://dx.doi.org/10.1007%2Fs11023-012-9282-2
  49. Bostrom, Nick (2014). "Chapter 10: Oracles, genies, sovereigns, tools (page 147)". Superintelligence: Paths, Dangers, Strategies. Oxford: Oxford University Press. ISBN 9780199678112. "For example, consider the risk that an oracle will answer questions not in a maximally truthful way but in such a way as to subtly manipulate us into promoting its own hidden agenda. One way to slightly mitigate this threat could be to create multiple oracles, each with a slightly different code and a slightly different information base. A simple mechanism could then compare the answers given by the different oracles and only present them for human viewing if all the answers agree." 
  50. "Intelligent Machines: Do we really need to fear AI?". BBC News. 27 September 2015. https://www.bbc.com/news/technology-32334568. 
  51. Marcus, Gary; Davis, Ernest (6 September 2019). "Opinion | How to Build Artificial Intelligence We Can Trust (Published 2019)". The New York Times. https://www.nytimes.com/2019/09/06/opinion/ai-explainability.html. 
  52. Sotala, Kaj; Yampolskiy, Roman (19 December 2014). "Responses to catastrophic AGI risk: a survey". Physica Scripta 90 (1): 018001. doi:10.1088/0031-8949/90/1/018001. Bibcode: 2015PhyS...90a8001S.  https://dx.doi.org/10.1088%2F0031-8949%2F90%2F1%2F018001
More
Upload a video for this entry
Information
Subjects: Others
Contributor MDPI registered users' name will be linked to their SciProfiles pages. To register with us, please refer to https://encyclopedia.pub/register :
View Times: 9.9K
Entry Collection: HandWiki
Revision: 1 time (View History)
Update Date: 22 Nov 2022
Notice
You are not a member of the advisory board for this topic. If you want to update advisory board member profile, please contact office@encyclopedia.pub.
OK
Confirm
Only members of the Encyclopedia advisory board for this topic are allowed to note entries. Would you like to become an advisory board member of the Encyclopedia?
Yes
No
${ textCharacter }/${ maxCharacter }
Submit
Cancel
There is no comment~
${ textCharacter }/${ maxCharacter }
Submit
Cancel
${ selectedItem.replyTextCharacter }/${ selectedItem.replyMaxCharacter }
Submit
Cancel
Confirm
Are you sure to Delete?
Yes No
Academic Video Service