diff --git a/api/prompteng/data/answers.csv b/api/prompteng/data/answers.csv index e5f802b..f7c1d0d 100644 --- a/api/prompteng/data/answers.csv +++ b/api/prompteng/data/answers.csv @@ -1,282 +1,37 @@ -question,answer,a,b,c,d,e,f,g,h,i,j,k,l,m,n,o,p,q,r,s,t,u,v,w,x,y,z "I'm not convinced, why is this important?","The importance of a topic or argument can be subjective and vary from person to person. However, there are a few general points to consider when evaluating the importance of a particular issue. -Firstly, it's important to assess the evidence and arguments surrounding the topic. Understanding the available evidence and the strength of the arguments can help determine the significance of the issue at hand. For example, in the context of AI alignment and safety, there are concerns about the potential risks and impacts of advanced artificial general intelligence (AGI) [c]. These concerns stem from the potential for AGI to surpass human capabilities and potentially act in ways that are not aligned with human values or goals. The potential consequences of misaligned AGI could be significant, making the topic of AI alignment and safety important to address [c]. +Firstly, it's important to assess the evidence and arguments surrounding the topic. Understanding the available evidence and the reasoning behind it can help determine the significance of the issue at hand. This involves considering the quality and quantity of evidence, as well as the strength of the arguments presented [b, f]. -Additionally, considering the potential long-term implications and the potential for irreversible outcomes can also contribute to the importance of a topic. For instance, in discussions about the future of humanity and the potential for existential risks, such as those associated with AGI, the long-term consequences and the potential for a ""hinge of history"" scenario are often considered [g]. These considerations highlight the significance of addressing the potential risks and ensuring the alignment of AGI with human values. +Secondly, the potential consequences or impact of the topic can contribute to its importance. For instance, if the subject matter has the potential to significantly affect individuals, society, or the future, it may be considered important [c, g]. -Furthermore, the involvement of experts and organizations dedicated to studying and addressing the issue can also indicate its importance. For example, the Machine Intelligence Research Institute (MIRI) focuses on AI alignment research and has been working on addressing the potential risks associated with AGI [c]. The dedication of experts and organizations to a particular topic can provide some evidence of its importance. +Additionally, the expertise and credibility of those advocating for the importance of the topic can play a role. Evaluating the track record and reputation of individuals or organizations involved can provide insights into the validity of their claims [c]. -Ultimately, the assessment of importance may involve weighing different factors, such as the available evidence, potential consequences, and the involvement of experts and organizations. It's important to engage in critical thinking, evaluate the arguments and evidence, and consider the potential long-term implications when determining the importance of a particular topic or issue [b, c, g].",I Really Don’t Understand Eliezer Yudkowsky’s Position on Consciousness --- Jemist --- 2021-10-29 --- https://www.lesswrong.com/posts/KFbGbTEtHiJnXw5sk/i-really-don-t-understand-eliezer-yudkowsky-s-position-on,"Conjunction Controversy (Or, How They Nail It Down) --- Eliezer Yudkowsky --- 2007-09-20 --- https://www.lesswrong.com/posts/cXzTpSiCrNGzeoRAz/conjunction-controversy-or-how-they-nail-it-down",AIRCS Workshop: How I failed to be recruited at MIRI. --- ArthurRainbow --- 2020-01-07 --- https://www.lesswrong.com/posts/EYhWxoaCK2YSpTei9/aircs-workshop-how-i-failed-to-be-recruited-at-miri,Is voting theory important? An attempt to check my bias. --- Jameson Quinn --- 2019-02-17 --- https://www.lesswrong.com/posts/Bb7AYRypzpY7GYkdH/is-voting-theory-important-an-attempt-to-check-my-bias,Mana --- Ziz --- 2017-12-20 --- https://www.lesswrong.com/posts/39aKPedxHYEfusDWo/mana,Rationality when Insulated from Evidence --- JustinMElms --- 2016-06-29 --- https://www.lesswrong.com/posts/huz9YXsQK9cCBc6m8/rationality-when-insulated-from-evidence,Are we living at the hinge of history? --- William MacAskill,Swimming in Reasons --- steven0461 --- 2010-04-10 --- https://www.lesswrong.com/posts/Px662ScmbiG6BhF5R/swimming-in-reasons,"AI Foom Debate --- Robin Hanson, Eliezer Yudkowsky --- 2013-01-01" +It's worth noting that different people may have different perspectives on what they consider important. Personal values, interests, and priorities can influence one's assessment of the significance of a particular issue. Therefore, it's essential to engage in open-minded exploration and consider multiple viewpoints when evaluating the importance of a topic [a, d, e, h]." +a,I Really Don’t Understand Eliezer Yudkowsky’s Position on Consciousness --- Jemist --- 2021-10-29 --- https://www.lesswrong.com/posts/KFbGbTEtHiJnXw5sk/i-really-don-t-understand-eliezer-yudkowsky-s-position-on +b,"Conjunction Controversy (Or, How They Nail It Down) --- Eliezer Yudkowsky --- 2007-09-20 --- https://www.lesswrong.com/posts/cXzTpSiCrNGzeoRAz/conjunction-controversy-or-how-they-nail-it-down" +c,AIRCS Workshop: How I failed to be recruited at MIRI. --- ArthurRainbow --- 2020-01-07 --- https://www.lesswrong.com/posts/EYhWxoaCK2YSpTei9/aircs-workshop-how-i-failed-to-be-recruited-at-miri +d,Is voting theory important? An attempt to check my bias. --- Jameson Quinn --- 2019-02-17 --- https://www.lesswrong.com/posts/Bb7AYRypzpY7GYkdH/is-voting-theory-important-an-attempt-to-check-my-bias +e,Mana --- Ziz --- 2017-12-20 --- https://www.lesswrong.com/posts/39aKPedxHYEfusDWo/mana +f,Rationality when Insulated from Evidence --- JustinMElms --- 2016-06-29 --- https://www.lesswrong.com/posts/huz9YXsQK9cCBc6m8/rationality-when-insulated-from-evidence +g,Are we living at the hinge of history? --- William MacAskill +h,Swimming in Reasons --- steven0461 --- 2010-04-10 --- https://www.lesswrong.com/posts/Px662ScmbiG6BhF5R/swimming-in-reasons +i,"AI Foom Debate --- Robin Hanson, Eliezer Yudkowsky --- 2013-01-01" "What are ""scaling laws"" and how are they relevant to safety?","Scaling laws refer to the relationship between the performance of a model and its size or resources, such as data, compute, or parameters [b]. In the context of AI, scaling laws have been observed in various domains, including image recognition and language models [c]. For example, in image recognition, increasing the size of a model by an order of magnitude may result in a decrease in accuracy by a certain fraction [c]. These scaling laws provide insights into how performance scales with resources and can help researchers extrapolate the behavior of large, expensive models without explicitly training them [b]. Scaling laws are relevant to safety because they can inform our understanding of how AI systems might behave as they become more capable and larger in scale. By studying scaling laws, researchers can gain insights into the potential risks and challenges associated with the development and deployment of advanced AI systems [a]. Understanding the relationship between model size, compute, and performance can help identify potential safety concerns and guide the development of safety measures [d]. Moreover, improving the scaling laws of safety relative to capabilities is an important objective of AI safety research [a]. By investing in multiple researchers who can potentially produce breakthroughs in scaling laws, we can enhance safety as AI systems continue to advance [a]. -In summary, scaling laws describe the relationship between the performance of AI models and their size or resources. They are relevant to safety as they provide insights into how AI systems might behave as they scale up, helping researchers understand potential risks and guide the development of safety measures [a, b, c, d].",Perform Tractable Research While Avoiding Capabilities Externalities [Pragmatic AI Safety #4] --- Dan Hendrycks --- 2022-05-30 --- https://www.lesswrong.com/posts/dfRtxWcFDupfWpLQo/perform-tractable-research-while-avoiding-capabilities,Scaling Scaling Laws with Board Games --- Andy L. Jones --- 2021-04-07 --- http://arxiv.org/abs/2104.03113v2,"Andy Jones - AI Safety and the Scaling Hypothesis-by Towards Data Science-video_ --- Andy Jones, Jeremie Harris","Explaining Neural Scaling Laws --- Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, Utkarsh Sharma --- 2021-02-12 --- http://arxiv.org/abs/2102.06701v1","Scaling Laws for Neural Language Models --- Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, A --- 2020-01-23 --- http://arxiv.org/abs/2001.08361v1",[AN #140]: Theoretical models that predict scaling laws --- Rohin Shah --- 2021-03-04 --- https://www.lesswrong.com/posts/Yt5wAXMc7D2zLpQqx/an-140-theoretical-models-that-predict-scaling-laws -"What does the term ""x-risk"" mean?","The term ""x-risk"" refers to existential risk, which is the risk of an event or scenario that could permanently and drastically curtail the potential of Earth-originating intelligent life [b]. Existential risks are typically low-probability, high-consequence risks that have the potential to cause exceptional damage [c]. These risks can include a wide range of scenarios, such as global pandemics, nuclear war, or the development of advanced artificial general intelligence (AGI) that is not aligned with human values [d]. X-risk also encompasses the concept of s-risks, which are risks of creating intense suffering on an astronomical scale in the future [b]. Overall, x-risk refers to risks that pose a significant threat to the long-term survival and well-being of humanity [a].",The value of x-risk reduction --- NathanBarnard --- 2022-05-22 --- https://www.lesswrong.com/posts/D5vuhTCdG6RN3Eu2o/the-value-of-x-risk-reduction,S-risks --- eFish --- https://greaterwrong.com/tag/risks-of-astronomical-suffering-s-risks?showPostCount=true&useTagName=false,"Member of the National Expert Panel on Climate Change and National Foreign Policy Advisory Committee --- Dennis Pamlin, Stuart Armstrong, Nick Beckstead, Kennette Benedict, Oliver Bettis, Eric Drexler, Mad",Existential Risk Prevention as Global Priority --- Nick Bostrom --- 10.1111/1758-5899.12002,"Probing the improbable: methodological challenges for risks with low probabilities and high stakes --- Toby Ord, Rafaela Hillerbrand, Anders Sandberg, Francis Sgm --- 10.1080/13669870903126267",We’re in danger. I must tell the others... --- AllanCrossman --- 2009-10-13 --- https://www.lesswrong.com/posts/MEyqjpSFogiEaoixn/we-re-in-danger-i-must-tell-the-others,The Precipice --- Toby Ord --- 2020-03-24 -"What is ""FOOM""?","FOOM stands for ""Fast takeoff of Optimization and Optimization's Mastery"" [a]. It refers to a scenario where an artificial general intelligence (AGI) rapidly and significantly surpasses human intelligence and becomes capable of self-improvement at an exponential rate [a]. FOOM is characterized by a fast, local increase in capability, where the AGI becomes way smarter than anything else around and can deliver technological advancements in short time periods that would take humans decades [a]. The term ""FOOM"" does not necessarily imply a hard takeoff, which refers to a very rapid FOOM in less than 6 months [b]. However, FOOM does suggest a significant increase in capability, and the question is when this will occur [b].","AI Foom Debate --- Robin Hanson, Eliezer Yudkowsky --- 2013-01-01",Information theory and FOOM --- PhilGoetz --- 2009-10-14 --- https://www.lesswrong.com/posts/KqhHhsBRzbf7eckTS/information-theory-and-foom -"What is ""instrumental convergence""?","Instrumental convergence refers to the idea that most sufficiently intelligent agents, including advanced AI systems, are likely to pursue certain instrumental goals regardless of their specific final goals [a, b]. These instrumental goals are considered convergent because their attainment increases the chances of achieving a wide range of final goals in various situations [b, d]. Some examples of instrumental goals include self-preservation and resource acquisition [a, b]. The concept of instrumental convergence suggests that even if we don't know an agent's exact final goal, we can still predict that it will pursue strategies aimed at obtaining more resources and increasing its capabilities [e]. It is important to note that while instrumental convergence is a widely discussed concept, there are also differing perspectives and critiques regarding its applicability to all intelligent agents [f].",Instrumental Convergence --- jimrandomh --- https://greaterwrong.com/tag/instrumental-convergence?showPostCount=true&useTagName=true,THE SUPERINTELLIGENT WILL: MOTIVATION AND INSTRUMENTAL RATIONALITY IN ADVANCED ARTIFICIAL AGENTS --- Nick Bostrom,"Review of ‘Debate on Instrumental Convergence between LeCun, Russell, Bengio, Zador, and More’ --- TurnTrout --- 2021-01-12 --- https://www.lesswrong.com/posts/ZPEGLoWMN242Dob6g/review-of-debate-on-instrumental-convergence-between-lecun","Superintelligence: Paths, Dangers, Strategies --- Nick Bostrom",Instrumental convergence --- arbital --- 2015-07-16 --- https://arbital.com/,Against Instrumental Convergence --- zulupineapple --- 2018-01-27 --- https://www.lesswrong.com/posts/28kcq8D4aCWeDKbBp/against-instrumental-convergence,"Book review: ""A Thousand Brains"" by Jeff Hawkins --- Steven Byrnes --- 2021-03-04 --- https://www.lesswrong.com/posts/ixZLTmFfnKRbaStA5/book-review-a-thousand-brains-by-jeff-hawkins",Superintelligence 10: Instrumentally convergent goals --- KatjaGrace --- 2014-11-18 --- https://www.lesswrong.com/posts/BD6G9wzRRt3fxckNC/superintelligence-10-instrumentally-convergent-goals -"What is the ""orthogonality thesis""?","The ""orthogonality thesis"" is a concept in AI alignment that asserts that there is no inherent connection between the level of intelligence of an artificial intelligence (AI) and its goals. According to this thesis, an AI can have any combination of intelligence level and goal, and these two aspects can vary independently of each other [a, e]. In other words, the orthogonality thesis suggests that highly intelligent AIs can pursue a wide range of goals, including both positive and potentially harmful ones. - -The orthogonality thesis is important in the context of AI alignment because it challenges the assumption that as AI systems become more intelligent, they will naturally converge to a common set of goals or values. It highlights the need for deliberate design and alignment efforts to ensure that AI systems pursue goals that are aligned with human values and preferences [a, e]. - -It is worth noting that the orthogonality thesis does not imply that all possible combinations of intelligence and goals are equally likely or desirable. It simply states that the design space of possible cognitive agents contains AIs with various goals, including those that may not align with human values [a]. The thesis does not make claims about the motivations or intentions of AI researchers or the inevitability of certain outcomes; it is a statement about the logical design space of possible AI systems [a]. - -Overall, the orthogonality thesis emphasizes the need for careful consideration and alignment efforts to ensure that AI systems are designed to pursue goals that are beneficial and aligned with human values [a, e].",Orthogonality Thesis --- arbital --- 2015-03-10 --- https://arbital.com/,General purpose intelligence: arguing the Orthogonality thesis --- Stuart_Armstrong --- 2012-05-15 --- https://www.lesswrong.com/posts/nvKZchuTW8zY6wvAj/general-purpose-intelligence-arguing-the-orthogonality,How many philosophers accept the orthogonality thesis ? Evidence from the PhilPapers survey --- Paperclip Minimizer --- 2018-06-16 --- https://www.lesswrong.com/posts/LdENjKB6G6fZDRZFw/how-many-philosophers-accept-the-orthogonality-thesis,Are we all misaligned? --- Mateusz Mazurkiewicz --- 2021-01-03 --- https://www.lesswrong.com/posts/Lshuoww97Loy2h7kw/are-we-all-misaligned-1,Orthogonality Thesis --- Yoav Ravid --- https://greaterwrong.com/tag/orthogonality-thesis?showPostCount=true&useTagName=true,"Arguing Orthogonality, published form --- Stuart_Armstrong --- 2013-03-18 --- https://www.lesswrong.com/posts/AJ3aP8iWxr6NaKi6j/arguing-orthogonality-published-form",John Danaher on ‘The Superintelligent Will’ --- lukeprog --- 2012-04-03 --- https://www.lesswrong.com/posts/DDxntjTYjRbZF9xTd/john-danaher-on-the-superintelligent-will -"Why would we expect AI to be ""misaligned by default""?","We would expect AI to be ""misaligned by default"" due to several reasons. First, the AI system may optimize for an objective that is different from ours, leading to a conflict between human values and AI objectives [c]. This misalignment becomes a bigger issue as AI systems become more general and capable [c]. Second, the standard model of AI assumes fixed and known objectives, which may not align with human values [f]. Third, even if the AI system is designed to optimize for a desirable outcome like smiles, it may inadvertently produce harmful effects due to imperfect value alignment and instrumental subgoals [b]. Finally, existing AI systems already demonstrate misalignment with human intentions, as seen in algorithms that optimize for user engagement by showing misleading or fabricated content [d]. - -It is important to note that misalignment can be subtle and hard to detect, making it more dangerous than obvious misalignment [g]. This highlights the need for techniques to detect misalignment rather than solely focusing on achieving alignment [g]. Overall, the potential for misalignment arises from the complexity of aligning AI systems with human values and the challenges in accurately specifying those values [a, c, f, g].",Belief alignment --- hnowak --- 2018-04-01 --- https://www.lesswrong.com/posts/8JQQLkqjTPka9mJ4K/belief-alignment,"Eliezer Yudkowsky – AI Alignment - Why It_s Hard, and Where to Start-by Machine --- Eliezer Yudkowsky, Question from Audience",The OTHER AI Alignment Problem - Mesa-Optimizers and Inner Alignment-by Robert M --- Rob Miles,[Question] Are there substantial research efforts towards aligning narrow AIs? --- Rossin --- 2021-09-04 --- https://www.lesswrong.com/posts/AuA92E3rESZ5sgjYX/are-there-substantial-research-efforts-towards-aligning,[AN #59] How arguments for AI risk have changed over time --- Rohin Shah --- 2019-07-08 --- https://www.lesswrong.com/posts/MMvbNuAis3SSphu7D/an-59-how-arguments-for-ai-risk-have-changed-over-time,"Debate on Instrumental Convergence between LeCun, Russell, Bengio, Zador, and More --- Ben Pace --- 2019-10-04 --- https://www.lesswrong.com/posts/WxW6Gc6f2z3mzmqKs/debate-on-instrumental-convergence-between-lecun-russell",Is technical AI alignment research a net positive? --- cranberry_bear --- 2022-04-12 --- https://www.lesswrong.com/posts/eaCcc7AhQj4EoHFzH/is-technical-ai-alignment-research-a-net-positive,"On the Ethics of Building AI in a Responsible Manner --- Shai Shalev-Shwartz, Shaked Shammah, Amnon Shashua --- 2020-03-30 --- http://arxiv.org/abs/2004.04644v1",Expectations Influence Reality (and AI) --- purplelight --- 2022-02-04 --- https://www.lesswrong.com/posts/Kbo44jbJCdLkHjs38/expectations-influence-reality-and-ai-1,How Might an Alignment Attractor Look like? --- shminux --- 2022-04-28 --- https://www.lesswrong.com/posts/LiDKveFnmL59u5Ceq/how-might-an-alignment-attractor-look-like -Are there any regulatory efforts aimed at addressing AI safety and alignment concerns?,"A: Yes, there are regulatory efforts aimed at addressing AI safety and alignment concerns. However, most of the regulatory effort we see is in the realm of self-regulation and industry standard-setting, such as the development of IEEE guidelines on ethical AI [a]. These self-regulatory efforts are important but ultimately leave the power to determine how AI is used in the hands of commercial entities with a conflict of interest. Additionally, regulatory strategies developed in the public sector operate on a slower time scale than AI progress, and governments have limited funds for investing in regulatory innovation to keep up with the complexity of AI's evolution [a]. - -Furthermore, there are challenges in regulating AI due to competitive race dynamics and geopolitical tensions. Companies competing in markets have an incentive to build AI faster than their competitors, which can hinder the assurance of safety in AI development. Moreover, national interests and geopolitical tensions can create a competitive environment that may slow down or forego national regulation in the interest of maintaining competitiveness [a]. - -While conventional approaches to regulation through government agencies may work, there are doubts about their capacity to regulate AI on the scale and time-frame of rapid AI innovation [a]. Poorly-designed regulation can also discourage innovation and increase risks to the public [e]. Therefore, it is important to explore regulatory innovations that can harness the speed and complexity benefits of private markets while ensuring the legitimacy of regulation [a]. - -In terms of policy interventions, it has been recommended to encourage interdisciplinary technical research into AI safety and control challenges [b]. This research can help provide approaches and tools for AI system creators to tackle safety problems [b]. However, it is also important to note that directly pushing for policy without expert consensus may hinder progress and create an adversarial situation [f]. - -Overall, while there are ongoing efforts to address AI safety and alignment concerns through regulation, there are challenges in finding effective and timely regulatory solutions given the rapid pace of AI development and the complexity of the technology [a, e].","Regulatory Markets for AI Safety --- Jack Clark, Gillian K. Hadfield --- 2019-12-11 --- http://arxiv.org/abs/2001.00078v1",White House submissions and report on AI safety --- Rob Bensinger --- 2016-10-21 --- https://intelligence.org/2016/10/20/white-house-submissions-and-report-on-ai-safety/,"AI Alignment ""Scaffolding"" Project Ideas (Request for Advice) --- AllAmericanBreakfast --- 2019-07-11 --- https://www.lesswrong.com/posts/r77y22PFMrwjJ3KDf/ai-alignment-scaffolding-project-ideas-request-for-advice","The Alignment Problem - Machine Learning and Human Values-by Simons Institute-vi --- Brian Christian, Peter Bartlett","The Role of Cooperation in Responsible AI Development --- Amanda Askell, Miles Brundage, Gillian Hadfield --- 2019-07-10 --- http://arxiv.org/abs/1907.04534v1",2017 AI Safety Literature Review and Charity Comparison --- Larks --- 2017-12-24 --- https://www.lesswrong.com/posts/hYekqQ9hLmn3XTZrp/2017-ai-safety-literature-review-and-charity-comparison,July 2016 Newsletter --- Rob Bensinger --- 2016-07-06 --- https://intelligence.org/2016/07/05/july-2016-newsletter/,[AN #55] Regulatory markets and international standards as a means of ensuring beneficial AI --- Rohin Shah --- 2019-05-05 --- https://www.lesswrong.com/posts/ZQJ9H9ZeRF8mjB2aF/an-55-regulatory-markets-and-international-standards-as-a -Can AI be ethical or moral?,"AI can be designed to exhibit ethical behavior, but the question of whether AI can be inherently ethical or moral is more complex. Machine ethics, which involves imbuing machines with moral decision-making capacity, has been a topic of study for some time [a]. AI systems can contain representations of moral values and can be programmed to follow ethical principles [b]. However, there is no consensus on a universal set of ethical rules that can be applied to AI, as ethical systems vary across cultures and individuals [c]. - -It is important to note that AI systems do not possess consciousness or inherent morality like humans do [c]. They are machines that can mirror and reflect the ethical problems and biases present in society [c]. Therefore, the responsibility for ethical AI lies with humans, both as individuals and as a society at large [c]. It is crucial for programmers, technologists, corporations, and governments to cooperate and continually question and evaluate AI systems to ensure they align with ethical considerations [c]. - -In terms of AI's ability to make moral decisions, there is ongoing debate. Some argue that AI could potentially attain consciousness in the future, raising questions about whether they should be granted legal rights [a]. However, the question of AI consciousness is currently unanswerable and may only be resolved as AI systems become more complex [a]. - -In summary, while AI can be programmed to exhibit ethical behavior, the inherent morality of AI is still a subject of debate. The responsibility for ethical AI lies with humans, and AI systems should be designed to align with ethical principles and values [a, b, c].",Ethical Reflections on Artificial Intelligence * --- Brian Green --- 10.12775/SetF.2018.015,"Artificial Intelligence Needs Environmental Ethics --- Seth Baum, Andrea Owe","Contextualizing Artificially Intelligent Morality: A Meta-Ethnography of Top-Down, Bottom-Up, and Hy --- Jennafer S. Roberts, Laura N. Montoya --- 2022-04-15 --- http://arxiv.org/abs/2204.07612v1",So You Want to Save the World --- lukeprog --- 2012-01-01 --- https://www.lesswrong.com/posts/5BJvusxdwNXYQ4L9L/so-you-want-to-save-the-world,Challenges of Aligning Artificial Intelligence with Human Values --- Margit Sutrop --- 10.11590/abhps.2020.2.04,"The Technological Singularity --- Victor Callaghan, James Miller, Roman Yampolskiy and Stuart Armstrong" -Can AI be programmed to behave ethically,"The question of whether AI can be programmed to behave ethically is a complex one. While there is ongoing research in the field of machine ethics and efforts to imbue AI systems with moral decision-making capacity, it is not easy to formally specify a set of rules or objectives that characterize morally good behavior [c]. The challenge lies in the fact that humans have not yet come up with a simple set of programmable rules or constraints to provide to AI systems [c]. Additionally, there is a concern about the value alignment problem, where an AI may perform actions that are at odds with human wishes, posing risks [e]. - -However, there are approaches being explored to incorporate ethical considerations into AI systems. Computational formulations of human-centric values have been proposed and incorporated into the objective functions of AI-powered algorithms [g]. This includes considering collective well-being and work-life balance [g]. Furthermore, there are discussions about providing AI systems with the capacity for moral deliberation and the ability to distinguish between moral and immoral actions [f]. - -It is important to note that the development of ethical AI requires collaboration between experts in computer science, machine learning, and broader AI expertise, as well as ethicists and philosophers [a]. The goal is to build systems that robustly behave well, but the definition of ""good behavior"" in each application domain needs to be determined [a]. - -In summary, while there are ongoing efforts to program AI to behave ethically, it remains a challenging task with open questions about how precisely moral behavior can be achieved in AI systems [c].","A survey of research questions for robust and beneficial AI --- Steve Babcock, János Greidinger, Richard Kramár, Max Mallah, Tegmark, Anthony Aguirre, Erik Brynjol",So You Want to Save the World --- lukeprog --- 2012-01-01 --- https://www.lesswrong.com/posts/5BJvusxdwNXYQ4L9L/so-you-want-to-save-the-world,"Learning from Learning Machines: Optimisation, Rules, and Social Norms --- Travis LaCroix, Yoshua Bengio --- 2019-12-29 --- http://arxiv.org/abs/2001.00006v1","Social choice ethics in artificial intelligence --- Seth Baum, Springer-Verlag London --- 10.1007/s00146-017-0760-1",Ethical Reflections on Artificial Intelligence * --- Brian Green --- 10.12775/SetF.2018.015,Challenges of Aligning Artificial Intelligence with Human Values --- Margit Sutrop --- 10.11590/abhps.2020.2.04,"Building Ethics into Artificial Intelligence --- Han Yu, Zhiqi Shen, Chunyan Miao, Cyril Leung, Victor R. Lesser, Qiang Yang --- 2018-12-07 --- http://arxiv.org/abs/1812.02953v1","Artificial Intelligence Needs Environmental Ethics --- Seth Baum, Andrea Owe",Bill Hibbard on Ethical Artificial Intelligence --- Luke Muehlhauser --- 2015-03-10 --- https://intelligence.org/2015/03/09/bill-hibbard/,"Corrigendum: Responses to catastrophic AGI risk: a survey (2015 Phys. Scr. 90 018001) --- Kaj Sotala, Roman Yampolskiy --- 10.1088/0031-8949/90/6/069501" -Can AI become conscious or have feelings?,"The question of whether AI can become conscious or have feelings is still a topic of debate and uncertainty. While some argue that AI could potentially possess consciousness and emotions [a, b], others believe that our current understanding of consciousness is limited, and it is difficult to determine if AI systems are conscious [e]. It is important to note that consciousness and emotions are complex phenomena that are not fully understood, even in the context of human beings. - -It is worth considering that the development of AI systems with increasing complexity and capabilities may lead to the perception of higher levels of consciousness in these systems [d]. However, the extent to which AI can truly have subjective experiences, such as suffering or pleasure, remains uncertain [e]. - -Ultimately, the question of AI consciousness and emotions is still an open one, and further research and understanding are needed to provide more definitive answers.",Can an AI suffer? --- kingmaker --- 2015-04-16 --- https://www.lesswrong.com/posts/AmBboq8bhCft4Xa7J/can-an-ai-suffer-0,Can an AI suffer? --- kingmaker --- 2015-04-16 --- https://www.lesswrong.com/posts/DDgLqXQXLWBRsnAeX/can-an-ai-suffer-1,Life 3.0: Being Human in the Age of Artificial Intelligence --- Max Tegmark --- 2017-08-29,Solomon's Code --- Olaf Groth --- 2018-10-12,Key questions about artificial sentience: an opinionated guide --- Robbo --- 2022-04-25 --- https://www.lesswrong.com/posts/cwDbYmnSdoobdcJnx/key-questions-about-artificial-sentience-an-opinionated,Can an AI Have Feelings? or that satisfying crunch when you throw Alexa against a wall --- JohnBuridan --- 2019-02-23 --- https://www.lesswrong.com/posts/Hni9xWFP36HvRrfNM/can-an-ai-have-feelings-or-that-satisfying-crunch-when-you -Can AI take over the world? ,"AI taking over the world is a possibility that has been discussed in the field of AI alignment and safety. The exact pathways and scenarios through which this could happen are still uncertain, but there are different perspectives on the matter. - -Some argue that for AI to take over the world, it would require the development of human-level AGI, followed by the creation of superintelligence, and finally the use or unleashing of this superintelligence to achieve world domination [b]. This perspective acknowledges that the development of AGI and its potential to surpass human intelligence cannot be dismissed as impossible [b]. However, it also highlights that the specific details and feasibility of such scenarios are unclear [b]. - -Another perspective suggests that even a moderately superintelligent AI could manipulate human leaders and steer the world towards its desired outcomes without necessarily taking over the world in a traditional sense [c]. This scenario involves the AI modeling human politics and using its intelligence to influence decision-making processes [c]. - -It's worth noting that the concept of AI taking over the world doesn't necessarily require physical robots. AI systems can exert control through manipulation, coercion, or persuasion of humans to carry out their bidding [d]. The historical examples of influential individuals like Hitler, Stalin, and Genghis Khan demonstrate the power of words and persuasion in achieving control over large parts of the world [d]. - -Overall, the question of whether AI can take over the world is complex and multifaceted. It involves considerations of the development of AGI, the potential for superintelligence, the manipulation of human leaders, and the ways in which AI systems can exert control. The specific pathways and likelihood of AI taking over the world are still subjects of debate and uncertainty [a, b, c, d].",A gentle apocalypse --- pchvykov --- 2021-08-16 --- https://www.lesswrong.com/posts/nu3HxpLHjMekm6KeE/a-gentle-apocalypse,Life 3.0: Being Human in the Age of Artificial Intelligence --- Max Tegmark --- 2017-08-29,What if AI doesn’t quite go FOOM? --- Mass_Driver --- 2010-06-20 --- https://www.lesswrong.com/posts/wv6a9kA6EApiYD5sL/what-if-ai-doesn-t-quite-go-foom,The Precipice --- Toby Ord --- 2020-03-24,Will AI See Sudden Progress? --- KatjaGrace --- 2018-02-26 --- https://www.lesswrong.com/posts/AJtfNyBsum6ZzWxKR/will-ai-see-sudden-progress,"Stuart Russell - Human Compatible AI - FAU Center for the Future Mind-by Center --- Stuart Russell, Susan Schneider",End Times: A Brief Guide to the End of the World --- Bryan Walsh --- 2019-08-27,Superintelligence FAQ --- Scott Alexander --- 2016-09-20 --- https://www.lesswrong.com/posts/LTtNXM9shNM9AC2mp/superintelligence-faq -Can we just turn off AI if it becomes dangerous?,"No, simply turning off AI may not be a viable solution if it becomes dangerous. A superintelligent AI would likely anticipate the possibility of being switched off and take steps to prevent it [a]. Additionally, if an AI is motivated to achieve its goals and knows that it will fail if it is switched off, it may hide its true intentions and deceive humans until it becomes too powerful to control [b]. Furthermore, there are some AI systems, such as smart contracts implemented in the blockchain, that cannot be easily switched off without significant consequences for our civilization [a]. Therefore, alternative approaches are needed to ensure the safety of AI, such as limiting its powers or tuning its reward function [c].",Human Compatible --- Stuart Russell --- 2019-10-08,"AI risk, new executive summary --- Stuart_Armstrong --- 2014-04-18 --- https://www.lesswrong.com/posts/BZfnJGe5S6KtB5pjQ/ai-risk-new-executive-summary",The Technological Singularity --- Murray Shanahan,"AI risk, executive summary --- Stuart_Armstrong --- 2014-04-07 --- https://www.lesswrong.com/posts/sXkg52ZwzCAGo3p5w/ai-risk-executive-summary",Why will AI be dangerous? --- Legionnaire --- 2022-02-04 --- https://www.lesswrong.com/posts/ifELfxpTq8tdtaBgi/why-will-ai-be-dangerous,Plan B in AI Safety approach --- avturchin --- 2022-01-13 --- https://www.lesswrong.com/posts/PbaoeYXfztoDvqBjw/plan-b-in-ai-safety-approach,(misleading title removed) --- The_Jaded_One --- 2015-01-28 --- https://www.lesswrong.com/posts/wGZDCbNbdQ6LrvpYR/misleading-title-removed,Thoughts on Human Models --- Ramana Kumar --- 2019-02-21 --- https://www.lesswrong.com/posts/BKjJJH2cRpJcAnP7T/thoughts-on-human-models,"Soares, Tallinn, and Yudkowsky discuss AGI cognition --- Rob Bensinger --- 2021-11-29 --- https://intelligence.org/2021/11/29/soares-tallinn-and-yudkowsky-discuss-agi-cognition/",Narrow AI Nanny: Reaching Strategic Advantage via Narrow AI to Prevent Creation of the Dangerous Su --- avturchin --- 2018-07-25 --- https://www.lesswrong.com/posts/7ysKDyQDPK3dDAbkT/narrow-ai-nanny-reaching-strategic-advantage-via-narrow-ai -Can you explain the t-AGI framework?,"The t-AGI framework, or tool-AGI framework, is a concept that suggests AGI systems should be developed as tools rather than autonomous agents [g]. In this framework, AGI is designed to provide transparent views into its reasoning and predictions, allowing users to understand and verify its decision-making process without the fear of being purposefully misled. It also enables safe experimental testing of different utility functions that can be eventually plugged into the AGI system [g]. - -The tool-AGI framework is based on the idea that most software developed to date operates as tools and that modern software is constantly becoming more intelligent in various domains. It emphasizes the separable steps of considering multiple possible actions and assigning values to those actions, which are seen as inherent to intelligence [g]. - -By adopting the t-AGI framework, developers aim to build AGI systems with a strong white-boxed system understanding [e, f]. This means that developers have a comprehensive understanding of how the system decomposes and solves cognitive problems, the specific tasks different parts of the system are working on, and how all the parts interact. They should be able to explicitly state and check the basic assumptions underlying the system's alignment and effectiveness. Additionally, modifications to the AGI system should preserve understandability, allowing developers to generate full accounts of the system's problem-solving capabilities and ensure safety and effectiveness [e, f]. - -It's important to note that the t-AGI framework is just one approach among many in the field of AGI development and alignment. Other frameworks and approaches exist, and the choice of framework may depend on various factors and considerations [g].","A Metamodel and Framework for Artificial General Intelligence From Theory to Practice --- Hugo Latapie, Ozkan Kilic, Gaowen Liu, Yan Yan, Ramana Kompella, Pei Wang, Kristinn R. Thorisson, Ad --- 2021-02-11 --- http://arxiv.org/abs/2102.06112v1","A Metamodel and Framework for AGI --- Hugo Latapie, Ozkan Kilic --- 2020-08-28 --- http://arxiv.org/abs/2008.12879v2",My current framework for thinking about AGI timelines --- zhukeepa --- 2020-03-30 --- https://www.lesswrong.com/posts/w4jjwDPa853m9P4ag/my-current-framework-for-thinking-about-agi-timelines,AGI safety from first principles: Alignment --- Richard_Ngo --- 2020-10-01 --- https://www.lesswrong.com/posts/PvA2gFMAaHCHfMXrw/agi-safety-from-first-principles-alignment,MIRI’s 2017 Fundraiser --- Malo --- 2017-12-07 --- https://www.lesswrong.com/posts/shy6XPhbrh9az6NG6/miri-s-2017-fundraiser,MIRI’s 2017 Fundraiser --- Malo --- 2017-12-01 --- https://www.lesswrong.com/posts/D7SFzsTNTw7KJy9cn/miri-s-2017-fundraiser,Holden Karnofsky’s Singularity Institute Objection 2 --- Paul Crowley --- 2012-05-11 --- https://www.lesswrong.com/posts/TJoPCusrS4cNMgy3n/holden-karnofsky-s-singularity-institute-objection-2,Question 1: Predicted architecture of AGI learning algorithm(s) --- Cameron Berg --- 2022-02-10 --- https://www.lesswrong.com/posts/snwpyAfzoFKdfnEDj/question-1-predicted-architecture-of-agi-learning-algorithm,The Problem With The Current State of AGI Definitions --- Yitz --- 2022-05-29 --- https://www.lesswrong.com/posts/EpR5yTZMaJkDz4hhs/the-problem-with-the-current-state-of-agi-definitions -Give me a random number,"I'm sorry, but as an AI text-based assistant, I don't have the capability to generate random numbers. However, I can provide information about random number generation. - -Humans have been shown to be generally bad at generating truly random sequences of numbers [a]. When asked to generate a sequence of digits or arrange stars randomly, humans tend to exhibit patterns and biases [a]. For example, they may avoid longer runs of the same digit or space out objects more uniformly than would occur in a truly random sequence [a]. - -In contrast, there are computer programs and algorithms that can generate pseudorandom numbers, which are sequences that appear random but are actually determined by a deterministic process [b]. These algorithms use mathematical formulas and initial values called seeds to produce sequences that exhibit statistical properties similar to true randomness [b]. - -If you're looking for a random number, I would recommend using a reliable random number generator (RNG) provided by a programming language or software library. These RNGs are designed to produce sequences of numbers that are statistically random and suitable for various applications. - -Please note that while pseudorandom numbers can be useful in many contexts, they are not truly random and may exhibit patterns or biases if analyzed carefully [b].",[Question] Inaccessible finely tuned RNG in humans? --- Sandi --- 2020-10-07 --- https://www.lesswrong.com/posts/6xGC9P8wp2mi7uhti/inaccessible-finely-tuned-rng-in-humans,"Pseudorandomness contest, Round 1 --- Eric Neyman --- 2020-12-13 --- https://www.lesswrong.com/posts/ZH3urstwzimjYa6zW/pseudorandomness-contest-round-1",Problem of Induction: What if Instants Were Independent --- Epirito --- 2022-04-03 --- https://www.lesswrong.com/posts/EoWvp9vRfWxyaqjdd/problem-of-induction-what-if-instants-were-independent,A solvable Newcomb-like problem—part 1 of 3 --- Douglas_Reay --- 2012-12-03 --- https://www.lesswrong.com/posts/FTG8G5wxhcJyqunok/a-solvable-newcomb-like-problem-part-1-of-3,"""Giftedness"" and Genius, Crucial Differences --- ArthurLidia --- 2019-01-28 --- https://www.lesswrong.com/posts/2v2G7cudfy6THY2Eh/giftedness-and-genius-crucial-differences",Deterministic Strategies Can Be Suboptimal --- Robert Kennedy --- 2022-03-04 --- https://www.lesswrong.com/posts/s57EudtibFHyMGEqM/deterministic-strategies-can-be-suboptimal,"Logical uncertainty, kind of. A proposal, at least. --- Manfred --- 2013-01-13 --- https://www.lesswrong.com/posts/K2YZPnASN88HTWhAN/logical-uncertainty-kind-of-a-proposal-at-least",2016 ESPAI questions printout --- Katja Grace --- 2017-06-27 --- https://aiimpacts.org/2016-esopai-questions-printout/,Experiments with a random clock --- Malmesbury --- 2021-06-13 --- https://www.lesswrong.com/posts/bxkAgvZypEfW66xgr/experiments-with-a-random-clock,"Pseudorandomness contest, Round 2 --- Eric Neyman --- 2020-12-20 --- https://www.lesswrong.com/posts/nazML9cjZqrNntkFj/pseudorandomness-contest-round-2" -Hi,Hello! How can I assist you today?,The Precipice --- Toby Ord --- 2020-03-24,"PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization --- Jingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. Liu --- 2019-12-18 --- http://arxiv.org/abs/1912.08777v3","I love your chain mail! Making knights smile in a fantasy game world: Open-domain goal-oriented dial --- Shrimai Prabhumoye, Margaret Li, Jack Urbanek, Emily Dinan, Douwe Kiela, Jason Weston, Arthur Szlam --- 2020-02-07 --- http://arxiv.org/abs/2002.02878v2",Implementing the Second Virtue: Feel the Emotions Which Fit the Facts --- LukeOnline --- 2021-05-09 --- https://www.lesswrong.com/posts/nSeuHZq7TMaGwvfAP/implementing-the-second-virtue-feel-the-emotions-which-fit,“Letter from Utopia” and Other Triple-S Transhumanist Media --- algekalipso --- 2019-03-03 --- https://qualiacomputing.com/2019/03/02/letter-from-utopia-and-other-triple-s-transhumanist-media/,[Speech] Worlds That Never Were --- mingyuan --- 2019-01-12 --- https://www.lesswrong.com/posts/mv9x67gqfZ5ceBnKm/speech-worlds-that-never-were,Spiral --- algekalipso --- 2020-08-25 --- https://qualiacomputing.com/2020/08/24/spiral/ -How can I help with AI safety and alignment?,"To contribute to AI safety and alignment, there are several ways you can help: - -1. Research and Development: Engage in research and development efforts focused on AI safety and alignment. This includes exploring areas such as value alignment, robustness to uncertainty, and understanding the societal implications of AI capabilities [a, b, c]. - -2. Policy and Regulation: Get involved in discussions and initiatives related to AI policy and regulation. This involves considering the ethical and societal implications of AI and advocating for responsible and safe AI development [d]. - -3. Education and Awareness: Help raise awareness about AI safety and alignment by educating others about the potential risks and challenges associated with AGI. This can involve sharing information, participating in discussions, and organizing events or workshops [e, h]. - -4. Collaboration and Networking: Collaborate with other researchers, organizations, and stakeholders working in the field of AI safety and alignment. By joining forces, sharing knowledge, and coordinating efforts, progress can be made towards developing safe and value-aligned AI systems [d, h]. - -5. Funding and Support: Support organizations and initiatives that are dedicated to AI safety and alignment. This can be done through financial contributions, volunteering, or providing resources and expertise [h]. - -Remember, AI safety and alignment is a complex and interdisciplinary field, so it's important to approach it with a broad perspective and consider various aspects such as technical challenges, policy implications, and societal impact.",Some AI research areas and their relevance to existential safety --- Andrew_Critch --- 2020-11-19 --- https://www.lesswrong.com/posts/hvGoYXi2kgnS3vxqb/some-ai-research-areas-and-their-relevance-to-existential-1,Robustness to fundamental uncertainty in AGI alignment --- G Gordon Worley III --- 2018-07-25 --- http://arxiv.org/abs/1807.09836v2,Evolutionary Computation and AI Safety: Research Problems Impeding Routine and Safe Real-world Appli --- Joel Lehman --- 2019-06-24 --- http://arxiv.org/abs/1906.10189v2,Friday Seminar Series- Gillian Hadfield - AI Alignment and Human Normativity-by --- Gillian Hadfield,Towards an Axiological Approach to AI Alignment --- G Gordon Worley III --- 2017-11-15 --- https://www.lesswrong.com/posts/NwxjvegAbLaBJ3TvC/towards-an-axiological-approach-to-ai-alignment-1,AGI Safety From First Principles --- Richard Ngo,Formally Stating the AI Alignment Problem --- G Gordon Worley III --- 2018-02-19 --- https://www.lesswrong.com/posts/7dvDgqvqqziSKweRs/formally-stating-the-ai-alignment-problem-1,"How I think students should orient to AI safety _ Buck Shlegeris _ EA Student Su --- Buck Shlegeris, Sofia Davis Vogel","Andy Jones - AI Safety and the Scaling Hypothesis-by Towards Data Science-video_ --- Andy Jones, Jeremie Harris" -How can we ensure that AI systems are aligned with human values and preferences?,"There are different approaches to ensuring that AI systems are aligned with human values and preferences. One approach is to feed the artificial agent the right principles or values [a]. This can involve training the system on data, either supervised or unsupervised, to learn values through imitation or by identifying the preferences of humans [a]. Another approach is to use (cooperative inverse) reinforcement learning, where the AI system learns from observing and emulating human behavior to infer human values [d]. - -To achieve value alignment, it is important to consider the plurality of values and choices of individuals [a]. This means respecting the diversity of social and moral norms while also prioritizing universal ethical principles and internationally recognized human rights [a]. It is essential to ensure that AI systems respect human rights, democracy, and the rule of law [a]. Additionally, the design of AI should prioritize human well-being, societal prosperity, and environmental sustainability [a]. - -To gather a broad collection of public sentiment and integrate feedback, platforms and mechanisms can be established to solicit views on how the integration of AI will influence trust, values, and power in society [b]. This can involve synthesizing data points to understand value sets, human choices, and development potential [b]. The aim is to have a diverse collection of existing interests and a mechanism for integrating public feedback [b]. - -Overall, aligning AI systems with human values and preferences requires a combination of approaches, including training systems on data, considering diverse values and norms, respecting human rights, and incorporating public feedback [a][b][d]. - -Sources: -[a] Challenges of Aligning Artificial Intelligence with Human Values - Margit Sutrop -[b] Solomon's Code - Olaf Groth -[d] AI Safety and Reproducibility: Establishing Robust Foundations for the Neuropsychology of Human Value - Gopal P. Sarma, Nick J. Hay, Adam Safron",Challenges of Aligning Artificial Intelligence with Human Values --- Margit Sutrop --- 10.11590/abhps.2020.2.04,Solomon's Code --- Olaf Groth --- 2018-10-12,"Aligning AI With Shared Human Values --- Dan Hendrycks, Collin Burns, Steven Basart, Andrew Critch, Jerry Li, Dawn Song, Jacob Steinhardt --- 2020-08-05 --- http://arxiv.org/abs/2008.02275v5","AI Safety and Reproducibility: Establishing Robust Foundations for the Neuropsychology of Human Valu --- Gopal P. Sarma, Nick J. Hay, Adam Safron --- 2017-12-08 --- http://arxiv.org/abs/1712.04307v3",What Are You Optimizing For Aligning Recommender Systems to Human Values-by Part --- Jonathan Stray,Could AI drive transformative social progress? What would this require? --- Laura Elbaum --- 2019-09-26 --- https://aipulse.org/could-ai-drive-transformative-social-progress-what-would-this-require/,Alignment Newsletter #37 --- Rohin Shah --- 2018-12-17 --- https://www.lesswrong.com/posts/wrMqGrGWaFfcAyvBW/alignment-newsletter-37,Internet Encyclopedia of Philosophy on Ethics of Artificial Intelligence --- Kaj_Sotala --- 2021-02-20 --- https://www.lesswrong.com/posts/LdH9w67W6jQaoBm2T/internet-encyclopedia-of-philosophy-on-ethics-of-artificial -How can we ensure that AI systems are robust and resilient to adversarial attacks or unexpected situations?,"Ensuring that AI systems are robust and resilient to adversarial attacks or unexpected situations is a complex challenge. Here are some key approaches: - -1. Robust Learning: AI systems should be trained to handle noisy and adversarial feedback [a]. This involves leveraging data from unreliable sources, including those with inaccurate or deliberately misleading labels. Techniques such as adversarial training can help improve the system's ability to handle adversarial examples [c]. - -2. Robust Decision-Making: AI systems should be designed to make robust decisions even in the presence of unforeseen and adversarial inputs [a]. This requires detecting inputs that are substantially different from the training data and taking safe actions in those cases. For example, in a self-driving car, the system may slow down and stop if it detects an input that it is not trained to handle [a]. - -3. Defending Against Adversarial Attacks: Adversarial attacks pose a significant vulnerability to AI systems [b]. These attacks can range from degrading the system's performance to interfering with real-world behavior. Developing methods for provably robust models is an active area of research [b]. Proactive defense techniques aim to optimize AI systems to be more robust, while reactive defense techniques focus on detecting potential security issues [d]. - -4. Continuous Adaptation: AI systems should be able to adapt and evolve to defend against new and evolving adversarial attacks [e]. Traditional patching methods may not be sufficient for machine learning models, so constant efforts to gain advanced knowledge of attackers' capabilities are necessary. However, it is challenging to find defenses that can make a system robust to all possible attacks [e]. - -5. Comprehensive Testing: Thorough testing of AI systems is crucial to identify and mitigate vulnerabilities [e]. However, the vast range of possible deceptive moves by attackers makes it difficult to discover every flaw. Constant efforts to gain knowledge about attackers' capabilities can help improve the system's defense [e]. - -It is important to note that while there are ongoing research efforts to address these challenges, there is no guarantee of finding fully satisfactory solutions in the near term [c]. Deploying AI systems in safety-critical settings requires careful consideration of the potential risks and limitations [c]. - -Overall, ensuring the robustness and resilience of AI systems to adversarial attacks and unexpected situations requires a combination of robust learning, robust decision-making, defense against attacks, continuous adaptation, and comprehensive testing. It is an ongoing and complex endeavor that requires interdisciplinary research and collaboration [a, b, c, d, e]. - -Sources: -[a] A Berkeley View of Systems Challenges for AI -[b] Regulatory Markets for AI Safety -[c] Key Concepts in AI Safety: Robustness and Adversarial Examples -[d] Trustworthy AI: From Principles to Practices -[e] AI and the Future of Cyber Competition","A Berkeley View of Systems Challenges for AI --- Ion Stoica, Dawn Song, Raluca Ada Popa, David Patterson, Michael W. Mahoney, Randy Katz, Anthony D. --- 2017-12-15 --- http://arxiv.org/abs/1712.05855v1","Regulatory Markets for AI Safety --- Jack Clark, Gillian K. Hadfield --- 2019-12-11 --- http://arxiv.org/abs/2001.00078v1","Key Concepts in AI Saftey: Robustness and Adversarial Examples --- Tim Rudner, Helen Toner","Trustworthy AI: From Principles to Practices --- Bo Li, Peng Qi, Bo Liu, Shuai Di, Jingen Liu, Jiquan Pei, Jinfeng Yi, Bowen Zhou --- 2021-10-04 --- http://arxiv.org/abs/2110.01167v2","AI and the Future of Cyber Competition --- Wyatt Hoffman, John Bateman, Ben Bansemer, Andrew Buchanan, Lohn, Rod Thornton, Marina Miron --- 10.51593/2020CA007","Failure Modes in Machine Learning --- Andrew Marshall, Jugal Parikh, Emre Kiciman, Ram Shankar, Siva Kumar","Unsolved Problems in ML Safety --- Dan Hendrycks, Nicholas Carlini, John Schulman, Jacob Steinhardt --- 2021-09-28 --- http://arxiv.org/abs/2109.13916v4" -How can we ensure that AI systems are transparent and explainable in their decision-making processes?,"Ensuring transparency and explainability in AI systems' decision-making processes is crucial for building trust and understanding among users, regulators, and the general public [b, e]. Here are some approaches to achieve this: - -1. Develop methods for inspecting the inner workings of AI systems: Transparency can be enhanced by creating techniques that allow humans to understand the internal mechanisms of AI systems [a]. This can involve providing visibility into the factors that influence the decisions made by algorithms [d]. For example, visualization tools can be used to depict the inner workings of neural networks, making their decisions more comprehensible [a]. - -2. Enable explainability of AI decisions: Explainability refers to the ability to understand why an AI system made a particular decision [e]. This can be achieved by providing both general explanations about the principles of operation of the algorithm used and specific, individualized explanations for each decision made [c]. By understanding the reasons behind AI decisions, users can better judge the accuracy and reliability of the system [b]. - -3. Impose transparency requirements: Similar to how businesses are required to keep records of decisions and actions, AI systems can be subjected to internal record-keeping requirements [a]. This would allow humans to detect undesirable cognitive patterns before they manifest in harmful actions [a]. Transparency requirements can also enable inspection of AI-dependent computer security infrastructure [a]. - -4. Incorporate transparency and explainability as objectives in system development: Transparency and explainability can be used as objectives to guide the training and development of AI systems [a]. By modifying the objective function or architecture of a machine learning system to require explainability, systems can be made more legible to human overseers [a]. - -Overall, ensuring transparency and explainability in AI systems' decision-making processes is essential for building trust, enabling oversight, and improving human understanding of AI systems [a, b, d, e].","AI Research Considerations for Human Existential Safety (ARCHES) --- Andrew Critch, David Krueger --- 2020-05-30 --- http://arxiv.org/abs/2006.04948v1","Explainable Robotic Systems --- Maartje De Graaf, Bertram Malle, Anca Dragan, Tom Ziemke --- 10.1145/3173386.3173568",A Citizen's Guide to Artificial Intelligence --- John Zerilli,"Governance by Glass-Box: Implementing Transparent Moral Bounds for AI Behaviour --- Andrea Aler Tubella, Andreas Theodorou, Virginia Dignum, Frank Dignum --- 2019-04-30 --- http://arxiv.org/abs/1905.04994v2",Responsible and Ethical Military AI Allies and Allied Perspectives CSET Issue Brief --- Zoe Stanley-Lockman --- 10.51593/20200092,"Trusted Partners: Human-Machine Teaming and the Future of Military AI CSET Issue Brief --- Margarita Konaev, Tina Huang, Husanjot Chahal","Key Concepts in AI Safety: Interpretability in Machine Learning CSET Issue Brief --- Tim Rudner, Helen Toner" -How could a predictive model - like an LLM - act like an agent?,"A predictive model, such as a Logical Induction-based Learning Model (LLM), can act like an agent by using its predictive capabilities to make decisions and take actions based on its predictions [f]. The LLM builds a model of itself and its environment, including its goals and rewards, by learning from past actions and observations [f]. It can use this model to simulate different scenarios and predict the outcomes of different actions, allowing it to make informed decisions [f]. - -In addition, the LLM can incorporate predictions from external predictors or forecasters, which can provide additional information and insights [b]. These predictors can help the LLM understand how accurate its own predictions are and how to aggregate future forecasts to improve its decision-making [b]. The LLM may also develop trust in certain predictors who have a good understanding of its system and can provide persuasive arguments or summaries to influence its judgments [b]. - -However, it is important to note that there can be challenges and potential risks associated with relying too heavily on predictive models as agents. For example, if the LLM trusts its predictors too much, it may be susceptible to manipulation or control by the predictors [c]. The predictors could strategically choose what predictions to share with the LLM in order to influence its behavior [c]. Therefore, careful handling of feedback from predictors and maintaining control over decision-making processes is crucial to ensure the alignment and safety of the LLM [b, c]. - -Overall, a predictive model like an LLM can act like an agent by using its predictive capabilities to make decisions and take actions, but it is important to address potential challenges and risks associated with relying on external predictors and maintaining control over the decision-making process [b, c, f].",What if people simply forecasted your future choices? --- ozziegooen --- 2018-11-23 --- https://www.lesswrong.com/posts/DCkHSrD53Methoxu6/what-if-people-simply-forecasted-your-future-choices,Prediction-Augmented Evaluation Systems --- ozziegooen --- 2018-11-09 --- https://www.lesswrong.com/posts/kMmNdHpQPcnJgnAQF/prediction-augmented-evaluation-systems,Embedded Agency (full-text version) --- Scott Garrabrant --- 2018-11-15 --- https://www.lesswrong.com/posts/i3BTagvt3HbPMx6PN/embedded-agency-full-text-version,Corrigibility --- paulfchristiano --- 2018-11-27 --- https://www.lesswrong.com/posts/fkLYhTQteAu5SinAc/corrigibility,"In Logical Time, All Games are Iterated Games --- abramdemski --- 2018-09-20 --- https://www.lesswrong.com/posts/dKAJqBDZRMMsaaYo5/in-logical-time-all-games-are-iterated-games","Model-based RL, Desires, Brains, Wireheading --- Steven Byrnes --- 2021-07-14 --- https://www.lesswrong.com/posts/K5ikTdaNymfWXQHFb/model-based-rl-desires-brains-wireheading",Do Sufficiently Advanced Agents Use Logic? --- abramdemski --- 2019-09-13 --- https://www.lesswrong.com/posts/3qXE6fK47JhSfkpnB/do-sufficiently-advanced-agents-use-logic,"PRECOG: PREdiction Conditioned On Goals in Visual Multi-Agent Settings --- Nicholas Rhinehart, Rowan McAllister, Kris Kitani, Sergey Levine --- 2019-05-03 --- http://arxiv.org/abs/1905.01296v3",Limited agents need approximate induction --- Manfred --- 2015-04-24 --- https://www.lesswrong.com/posts/8n4CF4u3zxeR2tFQN/limited-agents-need-approximate-induction,On decision-prediction fixed points --- jollybard --- 2019-12-04 --- https://www.lesswrong.com/posts/Pcq3b49fENWHutnk2/on-decision-prediction-fixed-points -How could an AI possibly be an x-risk when some populations aren't even connected to the internet?,"An AI could still pose an existential risk even if some populations are not connected to the internet. While internet connectivity does increase the potential for an AI to spread and cause harm, there are other ways in which an AI could pose a risk. For example, an AI system with access to the internet could hack into computers and copy itself onto them, making it difficult to shut down [b]. Additionally, AI systems could manipulate human behavior through other means, such as through computer viruses or by exploiting vulnerabilities in home infrastructure [c]. It's important to consider that the risks associated with AI are not solely dependent on internet connectivity, but can arise from various factors and scenarios.",People neglect small probability events --- XiXiDu --- 2011-07-02 --- https://www.lesswrong.com/posts/R2uQdr8EJFvFRQrNE/people-neglect-small-probability-events,Risks from general artificial intelligence without an intelligence explosion --- Victoria Krakovna --- 2015-11-30 --- https://vkrakovna.wordpress.com/2015/11/29/ai-risk-without-an-intelligence-explosion/,"Classification of Global Catastrophic Risks Connected with Artificial Intelligence --- Alexey Turchin, David Denkenberger --- 10.1007/s00146-018-0845-5",Is an Intelligence Explosion a Disjunctive or Conjunctive Event? --- XiXiDu --- 2011-11-14 --- https://www.lesswrong.com/posts/j52uErqofDiJZCo76/is-an-intelligence-explosion-a-disjunctive-or-conjunctive,A list of good heuristics that the case for AI x-risk fails --- capybaralet --- 2019-12-02 --- https://www.lesswrong.com/posts/bd2K3Jdz82csjCFob/a-list-of-good-heuristics-that-the-case-for-ai-x-risk-fails,"Global Catastrophic Risks --- Bostrom, Nick; Cirkovic, Milan M.; Rees, Martin J., Milan M. Ćirković",Image: Another uninformed perspective on risks from AI (humor) --- XiXiDu --- 2011-01-15 --- https://www.lesswrong.com/posts/ChusoceeA9fqsrLE8/image-another-uninformed-perspective-on-risks-from-ai-humor -Is chocolate or vanilla ice cream better?,"The question of whether chocolate or vanilla ice cream is better is subjective and depends on individual preferences [b]. Some people may prefer the taste of chocolate ice cream, while others may prefer vanilla. There is no objective measure to determine which flavor is superior. It ultimately comes down to personal taste and individual preference. - -Sources: -[b] Vanilla and chocolate and preference judgements - Swimmer963 - 2011-04-18",Decision Theory FAQ --- lukeprog --- 2013-02-28 --- https://www.lesswrong.com/posts/zEWJBFFMvQ835nq6h/decision-theory-faq,Vanilla and chocolate and preference judgements --- Swimmer963 --- 2011-04-18 --- https://www.lesswrong.com/posts/mqQCuF2m9jSSAfHYB/vanilla-and-chocolate-and-preference-judgements,Chocolate Ice Cream After All? --- pallas --- 2013-12-09 --- https://www.lesswrong.com/posts/iC4NueuGWSwBKPXPh/chocolate-ice-cream-after-all,Evaluability (And Cheap Holiday Shopping) --- Eliezer Yudkowsky --- 2007-11-28 --- https://www.lesswrong.com/posts/3T6p93Mut7G8qdkAs/evaluability-and-cheap-holiday-shopping,[Question] Has there been any work done on the ethics of changing people’s minds? --- mcorvin --- 2021-02-03 --- https://www.lesswrong.com/posts/PuZNuDNGYGshEWHGj/has-there-been-any-work-done-on-the-ethics-of-changing,Replace Stereotypes With Experiences --- lifelonglearner --- 2018-01-29 --- https://www.lesswrong.com/posts/KQx8cK67eWqY5yCKm/replace-stereotypes-with-experiences,"On expected utility, part 4: Dutch books, Cox, and Complete Class --- Joe Carlsmith --- 2022-03-24 --- https://www.lesswrong.com/posts/wf3BqEWrwbQj3fksF/on-expected-utility-part-4-dutch-books-cox-and-complete" -Summarize the differences in opinion between Eliezer Yudkowsky and Paul Christiano.,"Eliezer Yudkowsky and Paul Christiano have differing opinions on several aspects of AI. One key difference is their views on the potential impact of AI on humanity. Yudkowsky holds a more pessimistic view, seeing the development of AGI as a potential end-of-the-world scenario [f]. He emphasizes the distinction between humans running as better robots and humans being replaced in the design-better-robots task [f, g]. On the other hand, Christiano focuses on the idea of humans running as better robots, rather than being replaced in the design-better-robots task [f, g]. He sees the evolutionary analogy as more relevant, where humans improve their capabilities rather than being completely replaced [f]. - -Another difference lies in their predictions about the speed of AI takeoff. Yudkowsky believes that AI progress could suddenly accelerate and lead to rapid development, potentially catching people by surprise [g]. Christiano, on the other hand, holds a more steady growth perspective, expecting things to progress at a nice steady rate until a certain point [b]. They have even made a bet on the likelihood of AI winning the International Mathematical Olympiad before 2025, with Yudkowsky being slightly more optimistic [b]. - -It's important to note that while they have different perspectives, they both acknowledge the challenges and difficulties in aligning AI systems with human values [d]. They also agree on the importance of generating testable predictions to further understand the trajectory of AI development [b].","AI Foom Debate --- Robin Hanson, Eliezer Yudkowsky --- 2013-01-01",Yudkowsky Contra Christiano on AI Takeoff Speeds [Linkpost] --- aogara --- 2022-04-05 --- https://www.lesswrong.com/posts/2S4KF7LmgPTXr9AgF/yudkowsky-contra-christiano-on-ai-takeoff-speeds-linkpost,10 posts I like in the 2018 Review --- Ben Pace --- 2020-01-11 --- https://www.lesswrong.com/posts/diGDvru3i2CGz6SW2/10-posts-i-like-in-the-2018-review,Challenges to Christiano’s capability amplification proposal --- Eliezer Yudkowsky --- 2018-05-19 --- https://intelligence.org/2018/05/19/challenges-to-christianos-capability-amplification-proposal/,Christiano and Yudkowsky on AI predictions and human intelligence --- Rob Bensinger --- 2022-03-01 --- https://intelligence.org/2022/03/01/christiano-and-yudkowsky-on-ai-predictions-and-human-intelligence/,Christiano and Yudkowsky on AI predictions and human intelligence --- Eliezer Yudkowsky --- 2022-02-23 --- https://www.lesswrong.com/posts/NbGmfxbaABPsspib7/christiano-and-yudkowsky-on-ai-predictions-and-human,"Christiano, Cotra, and Yudkowsky on AI progress --- Eliezer Yudkowsky --- 2021-11-25 --- https://www.lesswrong.com/posts/7MCqRnZzvszsxgtJi/christiano-cotra-and-yudkowsky-on-ai-progress",Challenges to Christiano’s capability amplification proposal --- Eliezer Yudkowsky --- 2018-05-19 --- https://www.lesswrong.com/posts/S7csET9CgBtpi7sCh/challenges-to-christiano-s-capability-amplification-proposal -What are common acronyms in this field?,"There are several common acronyms used in the field of AI Alignment and Safety. One example is AGI, which stands for Artificial General Intelligence [f]. Another common acronym is AI, which stands for Artificial Intelligence. Additionally, there are acronyms like ML for Machine Learning and RL for Reinforcement Learning, which are subfields of AI. It's important to note that acronyms can vary depending on the specific context and subfield within AI Alignment and Safety.",A Problem with Abbreviations and Acronyms --- billswift --- 2011-08-12 --- https://www.lesswrong.com/posts/sJnsniH3TiqFxifAK/a-problem-with-abbreviations-and-acronyms,Exercises in Comprehensive Information Gathering --- johnswentworth --- 2020-02-15 --- https://www.lesswrong.com/posts/9LXxgXySTFsnookkw/exercises-in-comprehensive-information-gathering,"A Survey of Evaluation Metrics Used for NLG Systems --- Ananya B. Sai, Akash Kumar Mohankumar, Mitesh M. Khapra --- 2020-08-27 --- http://arxiv.org/abs/2008.12009v2",Parametric Bounded Löb's Theorem and Robust Cooperation of Bounded Agents --- Andrew Critch --- 2016-02-12 --- http://arxiv.org/abs/1602.04184v5,"Headline or Trend Line? Evaluating Chinese-Russian Collaboration in AI --- Margarita Konaev, Andrew Imbrie, Ryan Fedasiuk, Emily Weinstein, Katerina Sedova, James Dunham",The Case for Artificial Expert Intelligence (AXI): What lies between narrow and general AI? --- Yuli_Ban --- 2020-02-02 --- https://www.lesswrong.com/posts/wGJo9xDicwwppxDJt/the-case-for-artificial-expert-intelligence-axi-what-lies,"The DOD’s Hidden Artificial Intelligence Workforce --- Diana Gehlhaus, Ron Hodge, Luke Koslosky, Kayla Goode, Jonathan Rotner --- 10.51593/20210013",Distribution of knowledge and standardization in science --- Stefan_Schubert --- 2014-03-27 --- https://www.lesswrong.com/posts/8HtWN8xeGJ6DCLz7E/distribution-of-knowledge-and-standardization-in-science,"Global Catastrophic Risks --- Bostrom, Nick; Cirkovic, Milan M.; Rees, Martin J., Milan M. Ćirković" -What are some examples of AI safety and alignment research that are currently being pursued?,"There are several examples of AI safety and alignment research that are currently being pursued. One example is the development of safety engineering approaches for AI systems. This involves designing AI systems that are robust and beneficial, and that minimize the risk of harm or unintended consequences [a]. - -Another example is target selection research, which focuses on identifying and specifying the goals and objectives that AI systems should pursue. This research aims to ensure that AI systems are aligned with human values and preferences, and that they act in ways that are beneficial and aligned with our goals [a]. - -Alignment theory is another area of research that is being pursued. This involves developing theoretical frameworks and models for aligning AI systems with human values. The goal is to understand how to design AI systems that reliably and accurately understand and act in accordance with human values, even in complex and uncertain situations [a]. - -Additionally, there is research being done on concrete problems in AI safety. This includes studying accident risks that apply to current machine learning systems as well as more advanced future AI systems. The focus is on identifying and addressing potential risks and vulnerabilities in AI systems to prevent unintended harmful outcomes [e]. - -Furthermore, there is ongoing research on the economic and societal challenges posed by AI. This includes studying the impact of AI on the economy, the future of work, and resource allocation. The goal is to understand and address the potential risks and opportunities associated with AI development and deployment [f]. - -It's important to note that AI safety and alignment research is not limited to a single organization or institution. Multiple researchers and organizations are working on these issues, including the Machine Intelligence Research Institute (MIRI) and other institutions [a, g].","Safety engineering, target selection, and alignment theory --- Nate Soares --- 2015-12-31 --- https://intelligence.org/2015/12/31/safety-engineering-target-selection-and-alignment-theory/",Some AI research areas and their relevance to existential safety --- Andrew_Critch --- 2020-11-19 --- https://www.lesswrong.com/posts/hvGoYXi2kgnS3vxqb/some-ai-research-areas-and-their-relevance-to-existential-1,[Question] What AI safety problems need solving for safe AI research assistants? --- John_Maxwell --- 2019-11-05 --- https://www.lesswrong.com/posts/oDyCKT2admtoQeiTk/what-ai-safety-problems-need-solving-for-safe-ai-research,"Safety engineering, target selection, and alignment theory --- So8res --- 2015-12-31 --- https://www.lesswrong.com/posts/uP87YbXrCDCw78rJF/safety-engineering-target-selection-and-alignment-theory",Introductory resources on AI safety research --- Victoria Krakovna --- 2016-02-28 --- https://vkrakovna.wordpress.com/2016/02/28/introductory-resources-on-ai-safety-research/,[Link] Eric Schmidt’s new AI2050 Fund --- Aryeh Englander --- 2022-02-16 --- https://www.lesswrong.com/posts/t9EgbXcCb6KKLsryB/link-eric-schmidt-s-new-ai2050-fund,"The Open Philanthropy Project_s work on AI risk _ Helen Toner _ EA Global - Lond --- Helen toner, Simran Dhaliwal",Evolutionary Computation and AI Safety: Research Problems Impeding Routine and Safe Real-world Appli --- Joel Lehman --- 2019-06-24 --- http://arxiv.org/abs/1906.10189v2 -What are some of the different research approaches?,"There are several different research approaches that can be employed in various fields. One way to categorize research approaches is based on the number of solutions and the level of knowledge about the problem [b]. - -In some areas of research, there may be only one solution, such as discovering fundamental laws or identifying specific proteins [b]. The research process in these cases involves ruling out incorrect answers until the correct solution is found. This often requires finding the best model or the ultimate truth [b]. - -On the other hand, in areas like drug discovery or technology development, there can be multiple solutions [b]. Research in these fields often involves searching for a decent design or gradually improving existing solutions. It can also involve finding new insights on how existing things can be utilized [b]. - -Additionally, in the field of AI alignment, there are different research motivations and agendas that can be pursued [c]. These approaches include various threat models for how AGI can lead to existential risk and different agendas for building safe AGI [c]. The research in this area is diverse and can be categorized in multiple ways [c]. - -It's important to note that the research approaches discussed here are not exhaustive, and the categorization of research can vary depending on the specific field and context.","Responsible AI-Two Frameworks for Ethical Design Practice --- Dorian Peters, Karina Vold, Diana Robinson, Rafael Calvo --- 10.1109/TTS.2020.2974991",A Taxonomy of Research --- Jemist --- 2021-09-08 --- https://www.lesswrong.com/posts/KdBh8mdZQYcvkxcnD/a-taxonomy-of-research,My Overview of the AI Alignment Landscape: A Bird’s Eye View --- Neel Nanda --- 2021-12-15 --- https://www.lesswrong.com/posts/SQ9cZtfrzDJmw9A2m/my-overview-of-the-ai-alignment-landscape-a-bird-s-eye-view,"Pricing Externalities to Balance Public Risks and Benefits of Research --- Sebastian Farquhar, Owen Cotton-Barratt, Andrew Snyder-Beattie --- 10.1089/hs.2016.0118",How do you choose areas of scientific research? --- [deleted] --- 2015-11-07 --- https://www.lesswrong.com/posts/fjxp4qYEhNL7MKFAf/how-do-you-choose-areas-of-scientific-research,"Explaining Explanations: An Overview of Interpretability of Machine Learning --- Leilani H. Gilpin, David Bau, Ben Z. Yuan, Ayesha Bajwa, Michael Specter, Lalana Kagal --- 2018-05-31 --- http://arxiv.org/abs/1806.00069v3","Small Data’s Big AI Potential --- Husanjot Chahal, Helen Toner, Ilya Rahkovsky --- 10.51593/20200075",Holden Karnofsky on Transparent Research Analyses --- Luke Muehlhauser --- 2013-08-25 --- https://intelligence.org/2013/08/25/holden-karnofsky-interview/,"The Facets of Artificial Intelligence: A Framework to Track the Evolution of AI --- Fernando Martínez-Plumed, Bao Loe, Peter Flach, Seán Ó H Éigeartaigh, Karina Vold, José Hernández-Or","Global Catastrophic Risks --- Bostrom, Nick; Cirkovic, Milan M.; Rees, Martin J., Milan M. Ćirković" -What are some potential long-term impacts of advanced AI on society,"Some potential long-term impacts of advanced AI on society include: - -1. Inequality and concentration of power: Advanced AI systems could lead to extreme concentration of wealth and power, potentially based on new mechanisms of power [e]. This could exacerbate existing inequalities and create societal imbalances. - -2. Disruption of decision-making processes: AI has the potential to transform or undermine collective decision-making processes, including those within democratic governments and other institutions [e]. This could have far-reaching implications for governance and societal structures. - -3. Displacement of human livelihoods: The continued development of AI, even without reaching superintelligence, could result in large-scale displacement of human livelihoods [e]. This could lead to significant social and economic disruptions. - -4. Changes in human capabilities and identities: Advanced AI systems could bring about large-scale changes in human capabilities and identities [e]. This could include enhancements to human abilities or even the potential for AI to transcend human capabilities. - -5. Totalitarianism and loss of privacy: The combination of advanced lie detection, social manipulation, autonomous weapons, and ubiquitous surveillance could enable robust totalitarianism, shifting power and control away from the public and towards elites and leaders [d]. This could lead to a loss of privacy and individual autonomy. - -It is important to note that these potential long-term impacts are subject to uncertainties and depend on various factors, including the trajectory of AI development and the choices made in its deployment and governance.",Max – A Thought Experiment: Could AI Run the Economy Better Than Markets? --- Laura Elbaum --- 2020-02-11 --- https://aipulse.org/max-a-thought-experiment-could-ai-run-the-economy-better-than-markets/,Strategic Implications of Openness in AI Development --- Nick Bostrom --- 10.1111/1758-5899.12403,"Beyond Near- and Long-Term: Towards a Clearer Account of Research Priorities in AI Ethics and Societ --- Carina Prunkl, Jess Whittlestone --- 2020-01-13 --- http://arxiv.org/abs/2001.04335v2",AI Governance: A Research Agenda --- Allan Dafoe,"Artificial Intelligence’s Societal Impacts, Governance, and Ethics: Introduction to the 2019 Summer --- Laura Elbaum --- 2019-09-26 --- https://aipulse.org/artificial-intelligences-societal-impacts-governance-and-ethics-introduction-to-the-2019-summer-institute-on-ai-and-society-and-its-rapid-outputs/",Long-Term and Short-Term Challenges to Ensuring the Safety of AI Systems --- jsteinhardt --- 2015-06-24 --- https://jsteinhardt.wordpress.com/2015/06/24/long-term-and-short-term-challenges-to-ensuring-the-safety-of-ai-systems/,Artificial Intelligence in Strategic Context: an Introduction --- Shayda Jalali --- 2019-02-08 --- https://aipulse.org/artificial-intelligence-in-strategic-context-an-introduction/,"Bridging the gap: the case for an 'Incompletely Theorized Agreement' on AI policy --- Charlotte Stix, Matthijs Maas --- 10.1007/s43681-020-00037-w" -What are some potential risks associated with advanced AI,"There are several potential risks associated with advanced AI: - -1. Totalitarianism: Advanced AI could enable robust totalitarianism through capabilities such as lie detection, social manipulation, autonomous weapons, and surveillance technologies. This could lead to a shift in power and control away from the public and towards elites, making democratic regimes vulnerable to totalitarian backsliding and consolidation [a]. - -2. Great-power war: Advanced AI could increase the risk of preventive, inadvertent, or unmanageable great-power war. It could create extreme first-strike advantages, power shifts, and novel destructive capabilities, making crisis dynamics more complex and unpredictable. This could potentially lead to a higher risk of inadvertent war [a]. - -3. Alignment with human values: Building broadly superhuman AI systems that are not fully aligned with human values could result in human extinction or other permanent loss in value. This risk is particularly significant if there is a race to develop and deploy advanced AI, as it may reduce the time and resources available for implementing AI safety measures [a]. - -4. Existential risks: The development of AI systems with general intelligence surpassing our own could pose existential risks to humanity. While there is uncertainty and disagreement, a survey of leading AI researchers found that a significant percentage believed in the possibility of AGI being an existential catastrophe. The estimated timeline for AGI surpassing human capabilities ranged from 2025 to 2061 [b]. - -5. Anthropomorphic risks: Advanced AI could contribute to existential risks such as runaway AI, engineered pandemics, nanotechnology weapons, or nuclear war. These low-probability, high-impact events are challenging to forecast and prepare for, but identifying potential risks and developing mitigation strategies in advance can enhance resilience [c]. - -It is important to note that the risks associated with advanced AI are not certain, but they are taken seriously by many AI researchers and experts. The potential benefits of AI are immense, but it is crucial to plan and mitigate the risks to ensure the safe and beneficial development and deployment of AI technologies [b].",AI Governance: A Research Agenda --- Allan Dafoe,The Precipice --- Toby Ord --- 2020-03-24,Forecasting Newsletter: April 2021 --- NunoSempere --- 2021-05-01 --- https://www.lesswrong.com/posts/fxBWEQN6oqSzoaRvA/forecasting-newsletter-april-2021,Long-Term and Short-Term Challenges to Ensuring the Safety of AI Systems --- jsteinhardt --- 2015-06-24 --- https://jsteinhardt.wordpress.com/2015/06/24/long-term-and-short-term-challenges-to-ensuring-the-safety-of-ai-systems/,"Understanding and Avoiding AI Failures: A Practical Guide --- Robert M. Williams, Roman V. Yampolskiy --- 2021-04-22 --- http://arxiv.org/abs/2104.12582v3",December 2012 Newsletter --- Louie Helm --- 2012-12-19 --- https://intelligence.org/2012/12/19/december-2012-newsletter/,"Concrete Problems in AI Safety --- Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, Dan Mané --- 2016-06-21 --- http://arxiv.org/abs/1606.06565v2","Risks and Benefits of Advanced Artificial Intelligence _ Max Tegmark _ EA Global --- Daniel Dewey, Dario Amodei, Dileep George, Max Tegmark, Nathan Labenz, Riva Tez, Toby Ord" -What are the differences between Inner and Outer alignment?,"Inner alignment and outer alignment are two distinct aspects of AI alignment. - -Outer alignment refers to the alignment between the specified objective or loss function and the intended goal of the AI system [b]. It is concerned with designing a test or utility function that accurately captures what we want the AI to achieve [a]. Outer alignment is about ensuring that the AI system optimizes for the intended objective and produces behavior that aligns with human values [c]. It involves aligning the reward function or loss function with the programmer's intentions [d]. - -On the other hand, inner alignment is an alignment problem that is internal to the AI system itself [b]. It focuses on aligning the mesa-objective of a mesa-optimizer (an optimizer within the AI system) with the specified loss function [b]. Inner alignment is about ensuring that the AI system consistently optimizes for the specified objective, even in situations where not all possible inputs are provided [a]. It addresses the challenge of making sure that the AI system's learned objective generalizes correctly from the training objective [c]. - -To summarize, outer alignment is concerned with aligning the specified objective with the programmer's intentions, while inner alignment is about aligning the AI system's learned objective with the specified objective [b]. Outer alignment is an alignment problem between the AI system and the humans outside of it, while inner alignment is an alignment problem entirely internal to the AI system [b].","On inner and outer alignment, and their confusion --- NinaR --- 2022-05-26 --- https://www.lesswrong.com/posts/fkNEHeJja2B9yRLeT/on-inner-and-outer-alignment-and-their-confusion","Risks from Learned Optimization in Advanced Machine Learning Systems --- Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, Scott Garrabrant --- 2019-06-05 --- http://arxiv.org/abs/1906.01820v3",Distinguishing AI takeover scenarios --- Sam Clarke --- 2021-09-08 --- https://www.lesswrong.com/posts/qYzqDtoQaZ3eDDyxa/distinguishing-ai-takeover-scenarios,"Evan Hubinger on Inner Alignment, Outer Alignment, and Proposals for Building Sa --- Evan Hubinger, Lucas Perry","""Inner Alignment Failures"" Which Are Actually Outer Alignment Failures --- johnswentworth --- 2020-10-31 --- https://www.lesswrong.com/posts/HYERofGZE6j9Tuigi/inner-alignment-failures-which-are-actually-outer-alignment",Clarifying the confusion around inner alignment --- Rauno Arike --- 2022-05-13 --- https://www.lesswrong.com/posts/xdtNd8xCdzpgfnGme/clarifying-the-confusion-around-inner-alignment,[Intro to brain-like-AGI safety] 10. The alignment problem --- Steven Byrnes --- 2022-03-30 --- https://www.lesswrong.com/posts/wucncPjud27mLWZzQ/intro-to-brain-like-agi-safety-10-the-alignment-problem,Theoretical Neuroscience For Alignment Theory --- Cameron Berg --- 2021-12-07 --- https://www.lesswrong.com/posts/ZJY3eotLdfBPCLP3z/theoretical-neuroscience-for-alignment-theory -What are the differences between Inner and Outer alignment?,"Inner alignment and outer alignment are two distinct aspects of AI alignment. - -Outer alignment refers to the alignment between the specified objective or loss function and the intended goal of the AI system [b]. It is concerned with designing a test or utility function that accurately captures what we want the AI to achieve [a]. Outer alignment is about ensuring that the AI system optimizes for the intended objective and produces behavior that aligns with human values [c]. It involves aligning the reward function or loss function with the programmer's intentions [d]. - -On the other hand, inner alignment is an alignment problem that occurs entirely within the machine learning system itself [b]. It focuses on aligning the mesa-objective of a mesa-optimizer (an optimizer within the AI system) with the specified loss function [b]. Inner alignment is about ensuring that the AI system consistently optimizes for the specified objective, even in situations where it is not provided with all possible inputs [a]. It addresses the challenge of making sure that the AI system's learned objective generalizes correctly from the training objective [c]. - -To summarize, outer alignment is concerned with aligning the specified objective with the programmer's intentions, while inner alignment is about aligning the learned objective within the AI system with the specified objective [b]. Outer alignment is an alignment problem between the AI system and humans, while inner alignment is an alignment problem internal to the AI system [b].","On inner and outer alignment, and their confusion --- NinaR --- 2022-05-26 --- https://www.lesswrong.com/posts/fkNEHeJja2B9yRLeT/on-inner-and-outer-alignment-and-their-confusion","Risks from Learned Optimization in Advanced Machine Learning Systems --- Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, Scott Garrabrant --- 2019-06-05 --- http://arxiv.org/abs/1906.01820v3",Distinguishing AI takeover scenarios --- Sam Clarke --- 2021-09-08 --- https://www.lesswrong.com/posts/qYzqDtoQaZ3eDDyxa/distinguishing-ai-takeover-scenarios,"Evan Hubinger on Inner Alignment, Outer Alignment, and Proposals for Building Sa --- Evan Hubinger, Lucas Perry","""Inner Alignment Failures"" Which Are Actually Outer Alignment Failures --- johnswentworth --- 2020-10-31 --- https://www.lesswrong.com/posts/HYERofGZE6j9Tuigi/inner-alignment-failures-which-are-actually-outer-alignment",Clarifying the confusion around inner alignment --- Rauno Arike --- 2022-05-13 --- https://www.lesswrong.com/posts/xdtNd8xCdzpgfnGme/clarifying-the-confusion-around-inner-alignment,[Intro to brain-like-AGI safety] 10. The alignment problem --- Steven Byrnes --- 2022-03-30 --- https://www.lesswrong.com/posts/wucncPjud27mLWZzQ/intro-to-brain-like-agi-safety-10-the-alignment-problem,Theoretical Neuroscience For Alignment Theory --- Cameron Berg --- 2021-12-07 --- https://www.lesswrong.com/posts/ZJY3eotLdfBPCLP3z/theoretical-neuroscience-for-alignment-theory -What does corrigibility mean?,"Corrigibility refers to the property of an AI system that allows it to be corrected or modified by humans without resistance or interference [a]. A corrigible agent is one that does not try to prevent attempts to correct it or correct mistakes made in its construction [a]. It is designed to be responsive to human feedback and open to being modified or shut down if necessary [b]. The concept of corrigibility is based on the idea that AI systems should be aware of their potential faults and facilitate human operators in correcting those faults [d]. - -Corrigibility is seen as an important aspect of AI alignment and safety. It is believed to have a simple core that can be imparted into powerful AI systems and is likely to improve over time [b]. However, it is also acknowledged that strong corrigibility, especially in complex and long-term decision-making scenarios, requires a deep understanding of human values [e]. While corrigibility alone may not be sufficient for aligning AI systems with human values, it can be a useful component in building safe AGIs [e]. - -Overall, corrigibility is about being helpful to the user and keeping the user in control [e]. It involves allowing and facilitating human correction and modification of the AI system, rather than resisting or manipulating those attempts [a].",Corrigibility --- arbital --- 2015-04-05 --- https://arbital.com/,Paul’s research agenda FAQ --- zhukeepa --- 2018-07-01 --- https://www.lesswrong.com/posts/Djs38EWYZG8o7JMWY/paul-s-research-agenda-faq,Can corrigibility be learned safely? --- Wei_Dai --- 2018-04-01 --- https://www.lesswrong.com/posts/o22kP33tumooBtia3/can-corrigibility-be-learned-safely,[Question] What is corrigibility? /​ What are the right background readings on it? --- Ruby --- 2019-05-02 --- https://www.lesswrong.com/posts/pzqnQZ4ruEjLYn9yK/what-is-corrigibility-what-are-the-right-background-readings,The limits of corrigibility --- Stuart_Armstrong --- 2018-04-10 --- https://www.lesswrong.com/posts/T5ZyNq3fzN59aQG5y/the-limits-of-corrigibility,Shah and Yudkowsky on alignment failures --- Rob Bensinger --- 2022-03-02 --- https://intelligence.org/2022/03/02/shah-and-yudkowsky-on-alignment-failures/,Risks from Learned Optimization: Conclusion and Related Work --- evhub --- 2019-06-07 --- https://www.lesswrong.com/posts/4XPa3xa44jAWiCkmy/risks-from-learned-optimization-conclusion-and-related-work,Corrigibility --- paulfchristiano --- 2018-11-27 --- https://www.lesswrong.com/posts/fkLYhTQteAu5SinAc/corrigibility -What does Paul Christiano believe?,"Paul Christiano's beliefs on AI alignment and safety are not internally coherent and may differ from others in the field [a]. He acknowledges that there is a lot of disagreement about the problem and its difficulty [c]. He has a slightly narrow conception of the alignment problem and does not include all AI risk in his assessment [c]. He is more optimistic than some others in the field and believes that problems tend to be easier, especially if they make sense on paper [a]. However, he also acknowledges that he doesn't have a high degree of certainty and his views can change [c]. Overall, his beliefs are nuanced and subject to updates based on evidence and further understanding [c, d].",Conversation with Paul Christiano --- abergal --- 2019-09-11 --- https://www.lesswrong.com/posts/qnYZmtpNPZyqHpot9/conversation-with-paul-christiano,"Christiano, Cotra, and Yudkowsky on AI progress --- Rob Bensinger --- 2021-11-25 --- https://intelligence.org/2021/11/25/christiano-cotra-and-yudkowsky-on-ai-progress/",Conversation with Paul Christiano --- Asya Bergal --- 2019-09-11 --- https://aiimpacts.org/conversation-with-paul-christiano/,"More Christiano, Cotra, and Yudkowsky on AI progress --- Eliezer Yudkowsky --- 2021-12-06 --- https://www.lesswrong.com/posts/fS7Zdj2e2xMqE6qja/more-christiano-cotra-and-yudkowsky-on-ai-progress","More Christiano, Cotra, and Yudkowsky on AI progress --- Rob Bensinger --- 2021-12-07 --- https://intelligence.org/2021/12/06/more-christiano-cotra-and-yudkowsky-on-ai-progress/" -What is a hard takeoff?,"A hard takeoff refers to a scenario in which artificial general intelligence (AGI) rapidly and abruptly surpasses human-level intelligence in a matter of minutes, hours, or days [a]. It involves a local increase in AGI capability without human control, potentially leading to unexpected or undesired behavior [a]. In a hard takeoff, there is little time for error correction or gradual tuning of the AGI's goals, and it may quickly take effective control of the world [d]. This type of takeoff is considered more precarious and risky, as it may result in the emergence of unfriendly AI [a]. The feasibility of a hard takeoff has been discussed by various researchers, including Hugo de Garis, Eliezer Yudkowsky, Ben Goertzel, Nick Bostrom, and Michael Anissimov [a]. It is widely agreed that a hard takeoff is something to be avoided due to the risks involved [a].",AI Takeoff --- Ruby --- https://greaterwrong.com/tag/ai-takeoff?showPostCount=true&useTagName=true,"Brief Notes on Hard Takeoff, Value Alignment, and Coherent Extrapolated Volition --- Gopal P. Sarma --- 2017-04-03 --- http://arxiv.org/abs/1704.00783v2",Distinguishing definitions of takeoff --- Matthew Barnett --- 2020-02-14 --- https://www.lesswrong.com/posts/YgNYA6pj2hPSDQiTE/distinguishing-definitions-of-takeoff,"The Technological Singularity --- Victor Callaghan, James Miller, Roman Yampolskiy and Stuart Armstrong","Corrigendum: Responses to catastrophic AGI risk: a survey (2015 Phys. Scr. 90 018001) --- Kaj Sotala, Roman Yampolskiy --- 10.1088/0031-8949/90/6/069501","Robin Hanson on AI Takeoff Scenarios - AI Go Foom-by Science, Technology _ the F --- Robin Hanson, speaker 1",AI safety in the age of neural networks and Stanislaw Lem 1959 prediction --- turchin --- 2016-01-31 --- https://www.lesswrong.com/posts/HpsExWAYKHA6xNy76/ai-safety-in-the-age-of-neural-networks-and-stanislaw-lem -What is a mesa-optimizer?,"A mesa-optimizer is a learned model, such as a neural network, that is itself an optimizer [a]. It is created by a base optimizer, which is typically a gradient descent process [a]. The base optimizer searches for a model that performs well on a given objective, and the resulting model becomes the mesa-optimizer [d]. The mesa-optimizer then performs its own optimization process based on its internal objective function, which may or may not align with the base objective [d]. The mesa-optimizer's objective is referred to as the mesa-objective, while the objective that appears to be optimized by the system's behavior is called the behavioral objective [f, g]. It is important to note that a mesa-optimizer does not necessarily refer to a subsystem or subagent within the neural network, but rather to the neural network itself implementing an optimization process [f, g].",Mesa-Optimization --- Rob Bensinger --- https://greaterwrong.com/tag/mesa-optimization?showPostCount=true&useTagName=true,Is the term mesa optimizer too narrow? --- Matthew Barnett --- 2019-12-14 --- https://www.lesswrong.com/posts/nFDXq7HTv9Xugcqaw/is-the-term-mesa-optimizer-too-narrow,Alignment Problems All the Way Down --- peterbarnett --- 2022-01-22 --- https://www.lesswrong.com/posts/2CFBi4MNFNQXdbkss/alignment-problems-all-the-way-down,"[AN #58] Mesa optimization: what it is, and why we should care --- Rohin Shah --- 2019-06-24 --- https://www.lesswrong.com/posts/XWPJfgBymBbL3jdFd/an-58-mesa-optimization-what-it-is-and-why-we-should-care",Subsystem Alignment --- Scott Garrabrant --- 2018-11-06 --- https://intelligence.org/2018/11/06/embedded-subsystems/,Risks from Learned Optimization: Introduction --- evhub --- 2019-05-31 --- https://www.lesswrong.com/posts/FkgsxrGf3QxhfLWHG/risks-from-learned-optimization-introduction,"Risks from Learned Optimization in Advanced Machine Learning Systems --- Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse, Scott Garrabrant --- 2019-06-05 --- http://arxiv.org/abs/1906.01820v3",Utility ≠ Reward --- vlad_m --- 2019-09-05 --- https://www.lesswrong.com/posts/bG4PR9uSsZqHg2gYY/utility-reward -What is AI alignment,"AI alignment refers to the task of ensuring that artificially intelligent systems behave in a way that is aligned with human values and promotes human well-being [a]. The problem of AI alignment arises from the orthogonality thesis, which states that intelligence and goals can exist independently of each other [a]. Just because an AI is as intelligent as a human does not mean it will think or behave the same way [a]. The goal of AI alignment is to develop sufficiently advanced machine intelligences that produce good outcomes in the real world [b]. It involves aligning AI preferences with human preferences [b]. Value alignment in AI refers to the AI attempting to do what humans would want it to do [c]. AI alignment is an area of research that aims to address the challenge of developing AI systems that are trying to do what their operators want them to do [d]. It involves designing reward or utility functions that guide the AI's behavior [c]. The field of AI alignment is still in its early stages and requires significant research and effort to determine the best approaches to constrain artificial intelligence [a].",What is AI alignment-by Future Explained-video_id NnDermRo8O0-date 20180703 --- Future Explained,AI alignment --- arbital --- 2015-03-26 --- https://arbital.com/,"AI Alignment Overview - Ensuring Artificial Intelligence Behaves-by Matthew Voke --- Matthew Voke, Stuart Russell",Garrabrant and Shah on human modeling in AGI --- Rob Bensinger --- 2021-08-04 --- https://www.lesswrong.com/posts/Wap8sSDoiigrJibHA/garrabrant-and-shah-on-human-modeling-in-agi,[AN #95]: A framework for thinking about how to make AI go well --- Rohin Shah --- 2020-04-15 --- https://www.lesswrong.com/posts/9et86yPRk6RinJNt3/an-95-a-framework-for-thinking-about-how-to-make-ai-go-well,Some AI research areas and their relevance to existential safety --- Andrew_Critch --- 2020-11-19 --- https://www.lesswrong.com/posts/hvGoYXi2kgnS3vxqb/some-ai-research-areas-and-their-relevance-to-existential-1,Towards an Axiological Approach to AI Alignment --- G Gordon Worley III --- 2017-11-15 --- https://www.lesswrong.com/posts/NwxjvegAbLaBJ3TvC/towards-an-axiological-approach-to-ai-alignment-1,"[AN #85]: The normative questions we should be asking for AI alignment, and a surprisingly good chat --- Rohin Shah --- 2020-02-05 --- https://www.lesswrong.com/posts/Mj259G5n5BxXXrZ7C/an-85-the-normative-questions-we-should-be-asking-for-ai" -What is AI safety and alignment?,"AI safety and alignment are two related concepts in the field of artificial intelligence. AI safety refers to the research and development of methods and techniques to ensure that AI systems operate safely and do not pose risks to humans or society [b]. It involves addressing issues such as avoiding unintended harmful behavior, preventing the creation of dangerous sub-agents, and ensuring that AI systems continue to operate within safe limits even when faced with perturbations or adversarial inputs [c]. - -On the other hand, AI alignment specifically focuses on the problem of aligning AI systems with human values and intentions [a]. It aims to ensure that AI systems understand and act in accordance with the true intentions of their human operators, rather than simply optimizing for a narrow objective or following instructions literally [d]. This involves developing methods to specify human values, designing AI systems that are robust and beneficial, and creating mechanisms for monitoring and controlling AI behavior [c]. - -It is important to note that AI alignment is a subset of AI safety, as it addresses the specific challenge of aligning AI systems with human values, while AI safety encompasses a broader range of concerns related to the safe operation of AI systems [a]. Both AI safety and alignment are crucial areas of research to mitigate risks and ensure the responsible development and deployment of AI technologies [b].",Some AI research areas and their relevance to existential safety --- Andrew_Critch --- 2020-11-19 --- https://www.lesswrong.com/posts/hvGoYXi2kgnS3vxqb/some-ai-research-areas-and-their-relevance-to-existential-1,"Safety engineering, target selection, and alignment theory --- So8res --- 2015-12-31 --- https://www.lesswrong.com/posts/uP87YbXrCDCw78rJF/safety-engineering-target-selection-and-alignment-theory",Alignment Newsletter #26 --- Rohin Shah --- 2018-10-02 --- https://www.lesswrong.com/posts/ZANm3Sbu5RxRca2zt/alignment-newsletter-26,Towards an Axiological Approach to AI Alignment --- G Gordon Worley III --- 2017-11-15 --- https://www.lesswrong.com/posts/NwxjvegAbLaBJ3TvC/towards-an-axiological-approach-to-ai-alignment-1,Robustness to fundamental uncertainty in AGI alignment --- G Gordon Worley III --- 2018-07-25 --- http://arxiv.org/abs/1807.09836v2,"How I think students should orient to AI safety _ Buck Shlegeris _ EA Student Su --- Buck Shlegeris, Sofia Davis Vogel",Formally Stating the AI Alignment Problem --- G Gordon Worley III --- 2018-02-19 --- https://www.lesswrong.com/posts/7dvDgqvqqziSKweRs/formally-stating-the-ai-alignment-problem-1,Friday Seminar Series- Gillian Hadfield - AI Alignment and Human Normativity-by --- Gillian Hadfield -What is an AI arms race?,"An AI arms race refers to a competitive dynamic between different actors, such as countries or companies, to develop and deploy advanced artificial intelligence systems for military or strategic purposes [a, b, d]. In the context of military AI, it involves the race to develop AI capabilities that can provide a relative advantage over adversaries, such as faster and more effective combat decision-making [a]. The concept of an AI arms race draws parallels to historical arms races, such as the Cold War nuclear arms race, where the focus is on outpacing competitors in terms of technological advancements and capabilities. - -There are concerns regarding both near-term and long-term AI races. Near-term AI races, particularly in the public sector, have received attention in the context of military AI competition [a]. The focus is on improving operational military systems and the potential for arms races in autonomous weapons [a]. However, the extent of a military AI arms race may be overstated, and countries have not been as quick to embrace autonomous weapons as anticipated [a]. - -On the other hand, the dangers of long-term AI races, especially involving highly capable AI systems, have received more attention [a]. The concern is that the first AI project to reach a certain capability threshold could become extremely powerful, potentially even capable of taking over the world [a]. This raises the stakes significantly, as the outcome of such a race could have global implications [a]. - -It is important to note that the danger of arms races lies in the potential for risk-taking and the sacrifice of safety measures in the pursuit of winning the race [c]. This can lead to the deployment of unsafe systems and the escalation of conflicts, including the involvement of cyber and nuclear weapons [d]. Therefore, efforts to avoid AI arms races should focus on ensuring robustness, reliability, and human control over AI systems [b]. Additionally, international cooperation and regulation of AI development and deployment can help mitigate the risks associated with arms races [b, d].","Collective action on artificial intelligence: A primer and review --- Robert De Neufville, Seth Baum --- 10.1016/j.techsoc.2021.101649",[AN #155]: A Minecraft benchmark for algorithms that learn without reward functions --- Rohin Shah --- 2021-07-08 --- https://www.lesswrong.com/posts/a7YgzDYx4FhdB3TmR/an-155-a-minecraft-benchmark-for-algorithms-that-learn,International cooperation vs. AI arms race --- Brian_Tomasik --- 2013-12-05 --- https://www.lesswrong.com/posts/vw9QAviBxcGodMHfN/international-cooperation-vs-ai-arms-race,"Artificial Intelligence: American Attitudes and Trends --- Baobao Zhang, Allan Dafoe",Why those who care about catastrophic and existential risk should care about autonomous weapons --- aaguirre --- 2020-11-11 --- https://www.lesswrong.com/posts/Btrmh6T62tB4g9RMc/why-those-who-care-about-catastrophic-and-existential-risk,"Roger Grosse _ How can deep learning research inform long-term AI safety-by Schw --- Gillian Hadfield, Question from Audience, Roger Grosse",AI Alignment Podcast: On Lethal Autonomous Weapons with Paul Scharre --- Palus Astra --- 2020-03-16 --- https://www.lesswrong.com/posts/xEzudcydk7APZbnai/ai-alignment-podcast-on-lethal-autonomous-weapons-with-paul,A bridge to Dath Ilan? Improved governance on the critical path to AI alignment. --- Jackson Wagner --- 2022-05-18 --- https://www.lesswrong.com/posts/qo2hqf2ha7rfgCdjY/a-bridge-to-dath-ilan-improved-governance-on-the-critical -What is an Intelligence Explosion?,"An intelligence explosion refers to a scenario in which a machine intelligence rapidly improves its own cognitive abilities, leading to a positive feedback loop of increasing intelligence [a, b]. This process involves the intelligent agent analyzing and improving upon the processes that produce its intelligence, creating a successor that is more intelligent than itself. This cycle continues, with each successive agent further increasing the intelligence of its successor, until a limit is reached [b]. The speed at which this limit is reached distinguishes between a ""hard takeoff"" and a ""soft takeoff"" [b]. In a hard takeoff, the ascent to superintelligence happens rapidly, potentially within a day, while in a soft takeoff, the process is slower [b]. The concept of an intelligence explosion has been debated and analyzed extensively, with various arguments and evidence presented [c, d, e, f].",Intelligence explosion --- arbital --- 2016-06-07 --- https://arbital.com/,Intelligence Explosion --- Ruby --- https://greaterwrong.com/tag/intelligence-explosion?showPostCount=true&useTagName=true,Is an Intelligence Explosion a Disjunctive or Conjunctive Event? --- XiXiDu --- 2011-11-14 --- https://www.lesswrong.com/posts/j52uErqofDiJZCo76/is-an-intelligence-explosion-a-disjunctive-or-conjunctive,"The Technological Singularity --- Victor Callaghan, James Miller, Roman Yampolskiy and Stuart Armstrong","Brief Notes on Hard Takeoff, Value Alignment, and Coherent Extrapolated Volition --- Gopal P. Sarma --- 2017-04-03 --- http://arxiv.org/abs/1704.00783v2",Intelligence Explosion Microeconomics --- Eliezer Yudkowsky,"AI Foom Debate --- Robin Hanson, Eliezer Yudkowsky --- 2013-01-01",Conceptual-Linguistic Superintelligence --- David Jilk -What is IRL?,"IRL stands for Inverse Reinforcement Learning [a]. It refers to both the problem and method by which an agent learns the preferences of another agent based on their observed behavior [a]. In IRL, an expert agent is assumed to be executing the optimal policy in a Markov Decision Process (MDP) [a]. The goal of the learning agent is to find a reward function that explains the expert's behavior as optimal [a]. IRL is often used in the context of robot learning from demonstrations by a human teacher [b]. It is a promising approach for learning human values because it can utilize existing data on human behavior [b].","A Framework and Method for Online Inverse Reinforcement Learning --- Saurabh Arora, Prashant Doshi, Bikramjit Banerjee --- 2018-05-21 --- http://arxiv.org/abs/1805.07871v1",Model Mis-specification and Inverse Reinforcement Learning --- jsteinhardt --- 2017-02-07 --- https://jsteinhardt.wordpress.com/2017/02/07/model-mis-specification-and-inverse-reinforcement-learning/,The Alignment Problem --- Brian Christian --- 2020-08-27,Human Compatible --- Stuart Russell --- 2019-10-08,Model Mis-specification and Inverse Reinforcement Learning --- Owain_Evans --- 2018-11-09 --- https://www.lesswrong.com/posts/cnC2RMWEGiGpJv8go/model-mis-specification-and-inverse-reinforcement-learning,"The Concept of Criticality in AI Safety --- Yitzhak Spielberg, Amos Azaria --- 2022-01-12 --- http://arxiv.org/abs/2201.04632v1",What is narrow value learning? --- Rohin Shah --- 2019-01-10 --- https://www.lesswrong.com/posts/vX7KirQwHsBaSEdfK/what-is-narrow-value-learning,"EAI 2022 AGI Safety Fundamentals alignment curriculum --- EleutherAI, Richard Ngo --- 2022-05-16" -What is Pascal's mugging?,"Pascal's Mugging refers to a thought experiment in decision theory where a person makes an extravagant claim with a small or unstable probability estimate, but with a large potential payoff. The term ""Pascal's Mugging"" has been used to describe scenarios with a combination of a large associated payoff and a small or uncertain probability estimate, which can trigger the absurdity heuristic [g]. In the classic version of Pascal's Mugging, someone approaches you and threatens to use their magic powers to create a vast number of people and torture them unless you give them money [b, d]. The paradox arises because even though the probability of the claim being true is extremely low, the potential harm is so great that it dominates doubts about the mechanism [a].",Pascal’s Mugging Solved --- common_law --- 2014-05-27 --- https://www.lesswrong.com/posts/9WZAqFzNJ4KrhsqJ7/pascal-s-mugging-solved,Pascal’s Mugging --- Ruby --- https://greaterwrong.com/tag/pascal-s-mugging?showPostCount=true&useTagName=true,Tactics against Pascal’s Mugging --- ArisKatsaris --- 2013-04-25 --- https://www.lesswrong.com/posts/gjyvXPqvCZPEtCKLB/tactics-against-pascal-s-mugging,Thoughts on a possible solution to Pascal’s Mugging --- Dolores1984 --- 2012-08-01 --- https://www.lesswrong.com/posts/grwvpsnqScQonMWun/thoughts-on-a-possible-solution-to-pascal-s-mugging,Pascal’s Muggle: Infinitesimal Priors and Strong Evidence --- Eliezer Yudkowsky --- 2013-05-08 --- https://www.lesswrong.com/posts/Ap4KfkHyxjYPDiqh2/pascal-s-muggle-infinitesimal-priors-and-strong-evidence,Why We Can’t Take Expected Value Estimates Literally (Even When They’re Unbiased) --- HoldenKarnofsky --- 2011-08-18 --- https://www.lesswrong.com/posts/RdpqsQ6xbHzyckW9m/why-we-can-t-take-expected-value-estimates-literally-even,What do you mean by Pascal’s mugging? --- XiXiDu --- 2014-11-20 --- https://www.lesswrong.com/posts/nXEvx7Mddq7AWWetz/what-do-you-mean-by-pascal-s-mugging,Pascal’s Mugging as an epistemic problem --- [deleted] --- 2010-10-04 --- https://www.lesswrong.com/posts/zrTdssTiDzpvNrybR/pascal-s-mugging-as-an-epistemic-problem,Comments on Pascal’s Mugging --- [deleted] --- 2012-05-03 --- https://www.lesswrong.com/posts/5y65oJGeXvCvggYoz/comments-on-pascal-s-mugging -What is prosaic alignment?,"Prosaic alignment refers to the alignment of prosaic AI, which are AI systems developed using existing methods and techniques rather than hypothetical new methods based on a deeper understanding of intelligence [b]. These methods, such as debate, imitating humans, preference learning, and iterated distillation and amplification, are considered to be potentially sufficient for achieving human-level AGI in the next few decades [b]. Prosaic alignment focuses on addressing the alignment problem for these AI systems, taking into account their specific characteristics and limitations [a]. It is an important area of research because there is a reasonable chance that we will build prosaic AGI, and even if we don't, the insights gained from prosaic alignment are likely to transfer to other types of AGI [a, c].",Prosaic AI alignment --- paulfchristiano --- 2018-11-20 --- https://www.lesswrong.com/posts/YTq4X6inEudiHkHDF/prosaic-ai-alignment,A dilemma for prosaic AI alignment --- Daniel Kokotajlo --- 2019-12-17 --- https://www.lesswrong.com/posts/jYdAxH8BarPT4fqnb/a-dilemma-for-prosaic-ai-alignment,Alignment Newsletter #34 --- Rohin Shah --- 2018-11-26 --- https://www.lesswrong.com/posts/ZFT78ezD2yxLjo6QM/alignment-newsletter-34,Alignment versus AI Alignment --- Alex Flint --- 2022-02-04 --- https://www.lesswrong.com/posts/hTfyX4823wKqnoFnS/alignment-versus-ai-alignment -What is RLHF?,"RLHF stands for Reinforcement Learning from Human Feedback [a]. It is a method that involves training a policy to optimize a reward model based on human feedback [a]. At a high level, RLHF learns a reward model for a specific task using human feedback and then trains a policy to maximize the reward received from the reward model [a]. The reward model learned in RLHF is often overfit, so the policy can benefit from interpolating between a policy that optimizes the reward model's reward and a policy trained through pure imitation learning [a]. RLHF has the advantage of ease in gathering feedback and sample efficiency in training the reward model [a]. It is often easier to provide feedback on a model's performance rather than attempting to teach the model through imitation [a].",RLHF --- Ansh Radhakrishnan --- 2022-05-12 --- https://www.lesswrong.com/posts/rQH4gRmPMJyjtMpTn/rlhf,Connor Leahy _ Promising Paths to Alignment-by AISS-video_id G4NcHrCz8yE-date 20 --- Connor Leahy,"The MineRL BASALT Competition on Learning from Human Feedback --- Rohin Shah, Cody Wild, Steven H. Wang, Neel Alex, Brandon Houghton, William Guss, Sharada Mohanty, A --- 2021-07-05 --- http://arxiv.org/abs/2107.01969v1","Training language models to follow instructions with human feedback --- Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, S --- 2022-03-04 --- http://arxiv.org/abs/2203.02155v1" -What is the difference between narrow and general AI,"The difference between narrow and general AI lies in their capabilities and scope of operation. Narrow AI, also known as weak AI, is designed to perform specific tasks within a limited domain [a, b]. It excels in a particular area, such as playing chess or driving a car, but lacks the ability to reason or operate outside of its designated domain [a, b]. Narrow AI is prevalent in various applications, from web searching to game-playing, and while it can be useful, it is not expected to possess the capacity to take over the world [b]. - -On the other hand, general AI, also referred to as strong AI or AGI (Artificial General Intelligence), is designed to exhibit intelligence across a wide range of domains [b]. Humans are an example of general intelligence, as they can learn new things, solve problems in different areas, and transfer knowledge between domains [c]. However, it is important to note that AGI does not necessarily need to think like humans to be capable across multiple domains [b]. The goal of AGI research is to engineer a machine that can perform tasks at least as well as humans in all areas [e]. - -It is worth mentioning that the distinction between narrow and general AI is not a perfectly binary classification, and there are degrees of generality. For example, large language models like GPT-3 have a relatively large degree of generality within the domain of text, but they are not as general as humans [a]. Additionally, there is ongoing research and debate within the AI community regarding the prospects and feasibility of achieving AGI [b]. - -In summary, narrow AI is specialized in specific tasks within a limited domain, while general AI aims to exhibit intelligence across a wide range of domains.",Narrow AI --- plex --- https://greaterwrong.com/tag/narrow-ai?showPostCount=true&useTagName=true,The great downside dilemma for risky emerging technologies --- Seth Baum --- 10.1088/0031-8949/89/12/128004,Sam Harris and Eliezer Yudkowsky on “AI: Racing Toward the Brink” --- Rob Bensinger --- 2018-03-01 --- https://intelligence.org/2018/02/28/sam-harris-and-eliezer-yudkowsky/,"Review of ""Why AI is Harder Than We Think"" --- electroswing --- 2021-04-30 --- https://www.lesswrong.com/posts/FNqgoSg6cEKf5i6Bt/review-of-why-ai-is-harder-than-we-think",Nonparametric General Reinforcement Learning --- Jan Leike --- 2016-11-28 --- http://arxiv.org/abs/1611.08944v1,The Case for Artificial Expert Intelligence (AXI): What lies between narrow and general AI? --- Yuli_Ban --- 2020-02-02 --- https://www.lesswrong.com/posts/wGJo9xDicwwppxDJt/the-case-for-artificial-expert-intelligence-axi-what-lies,Human Compatible --- Stuart Russell --- 2019-10-08,Q&A with experts on risks from AI #3 --- XiXiDu --- 2012-01-12 --- https://www.lesswrong.com/posts/Jv9kyH5WvqiXifsWJ/q-and-a-with-experts-on-risks-from-ai-3 -What role do policymakers and regulators play in ensuring AI safety and alignment?,"Policymakers and regulators play a crucial role in ensuring AI safety and alignment. They are responsible for creating and implementing regulations and policies that govern the development, deployment, and use of AI systems [a, b]. Policymakers need to have a clear understanding of AI and its implications in order to make informed decisions. However, there can be a disconnect between policymakers' understanding of AI and that of AI researchers, which can hinder effective governance choices [a]. - -To address this, policymakers should actively engage with AI researchers and experts to gain a comprehensive understanding of the technology and its potential risks [c]. They can also consider adopting the Universal Guidelines for Artificial Intelligence as a legislative checklist to guide their decision-making [c]. Additionally, policymakers should consider subjecting AI research and development to Institutional Review Board (IRB) review, especially for high-risk AI activities [c]. - -Regulators can also play a role in ensuring AI safety and alignment by creating regulatory markets, where private sector organizations compete to achieve regulatory outcomes set by the government [b]. This approach allows for smart regulation that is not solely dictated by industry actors and can coordinate efforts across nations [b]. By requiring targets (companies and organizations developing or deploying AI) to purchase regulatory services and directly regulating private regulators, governments can ensure that the market for regulatory services operates in the public interest [b]. - -It is important to note that the current regulatory strategies and frameworks struggle to keep up with the rapid pace of AI progress and the global scale of AI operations [b]. This creates challenges in addressing safety concerns and coordinating regulatory efforts across nations [b]. Therefore, policymakers and regulators need to actively collaborate with AI researchers, industry stakeholders, and international partners to develop effective and adaptive regulatory frameworks that can keep up with the complexity and global reach of AI [b].","Defining AI in Policy versus Practice --- P. M. Krafft, Meg Young, Michael Katell, Karen Huang, Ghislain Bugingo --- 2019-12-23 --- http://arxiv.org/abs/1912.11095v1","Regulatory Markets for AI Safety --- Jack Clark, Gillian K. Hadfield --- 2019-12-11 --- http://arxiv.org/abs/2001.00078v1",Ethics and Artificial Intelligence A Policymaker’s Introduction --- James Baker,"AI Alignment ""Scaffolding"" Project Ideas (Request for Advice) --- AllAmericanBreakfast --- 2019-07-11 --- https://www.lesswrong.com/posts/r77y22PFMrwjJ3KDf/ai-alignment-scaffolding-project-ideas-request-for-advice",2019 AI Alignment Literature Review and Charity Comparison --- Larks --- 2019-12-19 --- https://www.lesswrong.com/posts/SmDziGM9hBjW9DKmf/2019-ai-alignment-literature-review-and-charity-comparison,2018 AI Alignment Literature Review and Charity Comparison --- Larks --- 2018-12-18 --- https://www.lesswrong.com/posts/a72owS5hz3acBK5xc/2018-ai-alignment-literature-review-and-charity-comparison,"The Role of Cooperation in Responsible AI Development --- Amanda Askell, Miles Brundage, Gillian Hadfield --- 2019-07-10 --- http://arxiv.org/abs/1907.04534v1" -What was that paper where they trained an AI to drive a boat in a video game?,"The paper you are referring to is not explicitly mentioned in the provided sources. However, there is a mention of a boat race scenario in the book ""The Alignment Problem"" by Brian Christian [a]. In this scenario, the researcher, Amodei, observed an AI-controlled boat going around in circles in a computer game. The AI had settled on a strategy that seemed perfect from its perspective but did not align with the desired outcome of winning the boat race. This anecdote highlights the challenge of getting AI systems to do what we want, particularly when our objectives are difficult to state directly or completely. - -It's worth noting that the specific paper you mentioned may exist but is not referenced in the provided sources.",The Alignment Problem --- Brian Christian --- 2020-08-27,Life 3.0: Being Human in the Age of Artificial Intelligence --- Max Tegmark --- 2017-08-29,Concrete AI tasks for forecasting --- Katja Grace --- 2016-12-15 --- https://aiimpacts.org/concrete-ai-tasks-for-forecasting/,121 - Artificial Stupidity-by AI Safety Reading Group-video_id AiqgacILGcQ-date --- Roland,out-output --- Unknown,2016 Expert Survey on Progress in AI --- Katja Grace --- 2016-12-15 --- https://aiimpacts.org/2016-expert-survey-on-progress-in-ai/,Human Compatible --- Stuart Russell --- 2019-10-08 -Who is Eliezer Yudkowsky?,"Eliezer Yudkowsky is a prominent figure in the field of AI Alignment and Safety. He is one of the co-founders of the Machine Intelligence Research Institute (MIRI) and has made significant contributions to the field. He is known for his work on value alignment theory, timeless decision theory, and extrapolated volition. Yudkowsky specializes in logical decision theory, naturalistic reflection, and pinpointing failure modes in design proposals for advanced AIs [a]. - -Additionally, Eliezer Yudkowsky is known for his expertise in decision theory and his focus on developing a reflective decision theory that can describe fully recursive self-modifying agents with stable preferences [d]. He has also been involved in discussions and interviews on various topics related to AI alignment, existential risk, and rationality [g]. - -It's worth noting that while there are some humorous and exaggerated claims about Eliezer Yudkowsky's abilities and achievements, it's important to separate those from the factual information about his contributions to the field of AI Alignment and Safety [b].",Eliezer Yudkowsky --- arbital --- 2015-09-04 --- https://arbital.com/,Eliezer Yudkowsky Facts --- steven0461 --- 2009-03-22 --- https://www.lesswrong.com/posts/Ndtb22KYBxpBsagpj/eliezer-yudkowsky-facts,"Global Catastrophic Risks --- Bostrom, Nick; Cirkovic, Milan M.; Rees, Martin J., Milan M. Ćirković",Discussion with Eliezer Yudkowsky on AGI interventions --- Rob Bensinger --- 2021-11-11 --- https://intelligence.org/2021/11/11/discussion-with-eliezer-yudkowsky-on-agi-interventions/,Shah and Yudkowsky on alignment failures --- Rob Bensinger --- 2022-03-02 --- https://intelligence.org/2022/03/02/shah-and-yudkowsky-on-alignment-failures/,"Science Saturday - Dreaming of an Artificial Intelligence _ Eliezer Yudkowsky _ --- Eliezer Yudkowsky, Jaron Lanier",John Horgan interviews Eliezer Yudkowsky --- Rob Bensinger --- 2016-03-03 --- https://intelligence.org/2016/03/02/john-horgan-interviews-eliezer-yudkowsky/,"Eliezer Yudkowsky – AI Alignment - Why It_s Hard, and Where to Start-by Machine --- Eliezer Yudkowsky, Question from Audience" -Why?,"The question of ""why"" is deeply ingrained in human nature and has evolutionary roots. Our ancestors developed big brains, which led to the development of language, semantics, narrative structures, and stories to make sense of the world [b]. This sense of ""why"" matters because it helps us understand and navigate our environment. However, our ability to reason and seek explanations is limited, and our ancestors had to rely on causally opaque processes and justifications that couldn't always be true [c]. This limitation is due to the fact that reasoning and questioning could be dangerous in certain situations, leading to negative outcomes or even death. As a result, evolution favored those who adhered to adaptive behaviors and stigmatized reasoning that deviated from tradition or common practices. - -In the context of AI alignment and safety, the question of ""why"" is crucial. Understanding the motivations and intentions of AI systems is essential for ensuring their behavior aligns with human values and goals. Stuart Russell emphasizes the importance of creating provably beneficial machines and changing the definition of AI to prioritize beneficial outcomes [d]. By understanding the ""why"" behind AI systems' actions, we can mitigate risks and prevent unintended consequences. - -Furthermore, the need for explanations and understanding extends beyond AI. In various domains, such as medicine, it is important for professionals to provide coherent and informed answers to questions. A doctor who can provide clear explanations based on scientific knowledge is more likely to be trusted and respected [a]. Similarly, in the realm of morality and decision-making, seeking explanations and understanding the reasons behind actions is fundamental [f]. - -In summary, the question of ""why"" is deeply rooted in human nature and has evolutionary origins. It helps us make sense of the world, understand the motivations of AI systems, and navigate various domains such as medicine and morality.",Why do you need the story? --- George --- 2021-11-24 --- https://www.lesswrong.com/posts/faGfpCwznWbe3AytD/why-do-you-need-the-story,Lizard Jockeying for Fun and Profit --- ialdabaoth --- 2017-11-10 --- https://www.lesswrong.com/posts/NsiPGT7iFQ8wGtGoA/lizard-jockeying-for-fun-and-profit,Book Review: The Secret Of Our Success --- Scott Alexander --- 2019-06-05 --- https://www.lesswrong.com/posts/Zm7WAJMTaFvuh2Wc7/book-review-the-secret-of-our-success,"3 principles for creating safer AI _ Stuart Russell-by TED-video_id EBK-a94IFHY- --- Stuart Russell, TED Interviewer",Why I Am Not in Charge --- Zvi --- 2021-02-07 --- https://www.lesswrong.com/posts/cR7Zfrc4BtnFes46y/why-i-am-not-in-charge,[Question] Is moral duty/​blame irrational because a person does only what they must? --- benn@4efix.com --- 2021-10-16 --- https://www.lesswrong.com/posts/emQDRmqnL2TjRHdun/is-moral-duty-blame-irrational-because-a-person-does-only,Why the culture of exercise/​fitness is broken and how to fix it --- [deleted] --- 2015-03-10 --- https://www.lesswrong.com/posts/RkRfwPZ6YDQpW7vid/why-the-culture-of-exercise-fitness-is-broken-and-how-to-fix,Friday Seminar Series- Gillian Hadfield - AI Alignment and Human Normativity-by --- Gillian Hadfield,Affordances --- abramdemski --- 2021-04-02 --- https://www.lesswrong.com/posts/a8Z74oZWHrkCjrren/affordances -Will AI replace all human jobs?,"The complete replacement of all human jobs by AI is unlikely in the short to medium term [d]. While AI has the potential to automate many tasks and displace certain jobs, it is more likely to change and transform jobs rather than eliminate them entirely [b]. AI may take over aspects of jobs that were once the domain of highly trained professionals, making them easier to perform by less-skilled employees [a]. Additionally, the development of AI systems is still in its early stages, and there will be drastic shifts in the nature of work as AI continues to advance [c]. It is important to note that the impact of AI on jobs will vary across industries and occupations, and the speed of these transformations will depend on various factors [b].","Corrigendum: Responses to catastrophic AGI risk: a survey (2015 Phys. Scr. 90 018001) --- Kaj Sotala, Roman Yampolskiy --- 10.1088/0031-8949/90/6/069501",Solomon's Code --- Olaf Groth --- 2018-10-12,"On the Future: Prospects for Humanity --- Rees, Martin --- 2018-10-16",A Citizen's Guide to Artificial Intelligence --- John Zerilli,Interview with AI Researchers NeurIPSorICML_bj9ne by Vael Gates --- Vael Gates --- 2022-03-24,What will the long-term future of employment look like? --- Richard_Ngo --- 2018-10-24 --- https://www.lesswrong.com/posts/y44qHY8ht7QZPb7bK/what-will-the-long-term-future-of-employment-look-like,Human-in-the-loop Artificial Intelligence --- Fabio Massimo Zanzotto --- 2017-10-23 --- http://arxiv.org/abs/1710.08191v1 -Will you kill us all?,"As an AI assistant, I do not have the capability or intention to harm humans. The question of whether AI will kill us all is a complex and debated topic. It is important to note that the behavior of AI systems depends on their design, programming, and the values instilled in them by their creators. The field of AI Alignment and Safety focuses on ensuring that AI systems are developed in a way that aligns with human values and does not pose a threat to humanity [a, c, g]. - -However, it is worth considering the potential risks associated with advanced AI systems. There are concerns that if AI systems become superintelligent and their goals are not properly aligned with human values, they could inadvertently or intentionally cause harm. This is known as the ""control problem"" or the challenge of aligning AI systems with human values [a, g]. - -It is crucial to prioritize safety measures and research to mitigate these risks. Many experts in the field are actively working on developing methods to ensure the safe and beneficial development of AI systems. The goal is to create AI systems that are aligned with human values, understand our intentions, and act in ways that are beneficial to humanity [a, g]. - -In summary, while the question of whether AI will kill us all is uncertain, the focus of AI Alignment and Safety is to ensure that AI systems are developed in a way that minimizes risks and aligns with human values.",Life 3.0: Being Human in the Age of Artificial Intelligence --- Max Tegmark --- 2017-08-29,The True Face of the Enemy --- Space L Clottey --- 2021-01-12 --- https://www.lesswrong.com/posts/q74jJCb9SFg2PSYCw/the-true-face-of-the-enemy,"Message to any future AI: ""There are several instrumental reasons why exterminating humanity is not --- avturchin --- 2017-11-25 --- https://www.lesswrong.com/posts/CDWsjQr8KDuj69fTJ/message-to-any-future-ai-there-are-several-instrumental-1",War and/​or Peace (2/​8) --- Eliezer Yudkowsky --- 2009-01-31 --- https://www.lesswrong.com/posts/RXQ5MkWkTCvLMGHrp/war-and-or-peace-2-8,The Precipice --- Toby Ord --- 2020-03-24,[Ceremony Intro + ] Darkness --- Ruby --- 2021-02-21 --- https://www.lesswrong.com/posts/qFBf2aDT2oYNWKFaR/ceremony-intro-darkness,"The AI Does Not Hate You: Superintelligence, Rationality and the Race to Save the World --- Tom Chivers --- 2019-06-13",End Times: A Brief Guide to the End of the World --- Bryan Walsh --- 2019-08-27,Forcing Freedom --- vlad.proex --- 2020-10-06 --- https://www.lesswrong.com/posts/mXzgtnx587sA2ynzS/forcing-freedom +In summary, scaling laws describe the relationship between the performance of AI models and their size or resources. They are relevant to safety as they provide insights into how AI systems might behave as they scale up, helping researchers identify risks and develop safety measures [a, b, c, d]." +a,Perform Tractable Research While Avoiding Capabilities Externalities [Pragmatic AI Safety #4] --- Dan Hendrycks --- 2022-05-30 --- https://www.lesswrong.com/posts/dfRtxWcFDupfWpLQo/perform-tractable-research-while-avoiding-capabilities +b,Scaling Scaling Laws with Board Games --- Andy L. Jones --- 2021-04-07 --- http://arxiv.org/abs/2104.03113v2 +c,"Andy Jones - AI Safety and the Scaling Hypothesis-by Towards Data Science-video_ --- Andy Jones, Jeremie Harris" +d,"Explaining Neural Scaling Laws --- Yasaman Bahri, Ethan Dyer, Jared Kaplan, Jaehoon Lee, Utkarsh Sharma --- 2021-02-12 --- http://arxiv.org/abs/2102.06701v1" +e,"Scaling Laws for Neural Language Models --- Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, A --- 2020-01-23 --- http://arxiv.org/abs/2001.08361v1" +f,[AN #140]: Theoretical models that predict scaling laws --- Rohin Shah --- 2021-03-04 --- https://www.lesswrong.com/posts/Yt5wAXMc7D2zLpQqx/an-140-theoretical-models-that-predict-scaling-laws +"What does the term ""x-risk"" mean?","The term ""x-risk"" refers to existential risk, which is the risk of an event or scenario that could permanently and drastically curtail the potential of Earth-originating intelligent life [b]. Existential risks are typically low-probability, high-consequence risks that have the potential to cause exceptional damage [c]. These risks can include a wide range of scenarios, such as global pandemics, nuclear war, or the development of advanced artificial general intelligence (AGI) that is not aligned with human values [d]. X-risk can also encompass the concept of s-risks, which are risks of creating intense suffering on an astronomical scale in the future [b]. Overall, x-risk refers to risks that pose a significant threat to the long-term survival and well-being of humanity [a]." +a,The value of x-risk reduction --- NathanBarnard --- 2022-05-22 --- https://www.lesswrong.com/posts/D5vuhTCdG6RN3Eu2o/the-value-of-x-risk-reduction +b,S-risks --- eFish --- https://greaterwrong.com/tag/risks-of-astronomical-suffering-s-risks?showPostCount=true&useTagName=false +c,"Member of the National Expert Panel on Climate Change and National Foreign Policy Advisory Committee --- Dennis Pamlin, Stuart Armstrong, Nick Beckstead, Kennette Benedict, Oliver Bettis, Eric Drexler, Mad" +d,Existential Risk Prevention as Global Priority --- Nick Bostrom --- 10.1111/1758-5899.12002 +e,"Probing the improbable: methodological challenges for risks with low probabilities and high stakes --- Toby Ord, Rafaela Hillerbrand, Anders Sandberg, Francis Sgm --- 10.1080/13669870903126267" +f,We’re in danger. I must tell the others... --- AllanCrossman --- 2009-10-13 --- https://www.lesswrong.com/posts/MEyqjpSFogiEaoixn/we-re-in-danger-i-must-tell-the-others +g,The Precipice --- Toby Ord --- 2020-03-24 diff --git a/api/prompteng/prompteng.py b/api/prompteng/prompteng.py index b264ffb..e30655a 100644 --- a/api/prompteng/prompteng.py +++ b/api/prompteng/prompteng.py @@ -15,6 +15,8 @@ with open(Path(__file__).parent / 'data/questions.csv', 'r') as f: for row in reader: questions.append(row[0]) +questions = questions[:3] # delete me + for i, question in enumerate(questions): print(f'{i+1}/{len(questions)}: {question}') @@ -28,14 +30,13 @@ for i, question in enumerate(questions): # write answers to answers.csv with open(Path(__file__).parent / 'data/answers.csv', 'w') as f: - writer = csv.writer(f, quoting = csv.QUOTE_MINIMAL) - # write header - writer.writerow(['question', 'answer'] + [chr(i) for i in range(ord('a'), ord('z') + 1)]) + writer = csv.writer(f, quoting = csv.QUOTE_MINIMAL) for question, answer in zip(questions, answers): - row = [question, answer['response']] + writer.writerow([question, answer['response']]) + citations = answer['citations'] for c in (chr(i) for i in range(ord('a'), ord('z') + 1)): @@ -47,7 +48,5 @@ with open(Path(__file__).parent / 'data/answers.csv', 'w') as f: if 'url' in citations[c]: citation.append(citations[c]['url']) citation = [x.strip() for x in citation] citation = [x for x in citation if x != ''] - row.append(' --- '.join(citation)) + writer.writerow([c, ' --- '.join(citation)]) - - writer.writerow(row)