top of page

SI WHITE PAPER 5: SAFE PERSONALIZED SUPERINTELLIGENCE

ABSTRACT: Safe Personalized SuperIntelligence

Personalized SuperIntelligence (PSI) represents the next leap forward in developing advanced, autonomous, artificial intelligence agents. However, because of their extreme intelligence, PSIs also represent a dangerous potential threat to human safety. This white paper describes how to design and construct such AI agents safely. It also shows how to use them as part of a safe, larger SuperIntelligent system.

​

Preferred implementations, including methods that rapidly enable PSIs to improve themselves, with or without human oversight, are described. The white paper describes how to produce different versions of PSIs using several novel methods. Methods for implementing scalable safety checks that remain effective even as PSIs become much more intelligent than their human creators are also covered.

​

Finally, an entirely new approach to AI safety is presented, combining a community of PSIs with proven blockchain methods. Rather than relying on testing to ensure safety, it envisions systems in which PSI safety is built in.

SUMMARY: Safe Personalized SuperIntelligence

This white paper describes how to build a Personalized SuperIntelligence, an AI agent customized to one person, and how to keep such agents safe once they exceed their owners in intelligence. The paper builds on Kaplan’s four earlier white papers in this series, which described how individual AI agents can be customized and put to work, how they apply to real products and platforms, how they can be integrated into a Human-Centered AGI network, and how safety and ethical information from many agents can be combined into a representative sample of human values. This paper carries that work forward, adding methods that can be used with the earlier systems or on their own.

​

What a Personalized SuperIntelligence is

A PSI begins as software customized, trained, and personalized to an individual owner. Because it improves itself, it becomes exponentially more intelligent than the person who created it. It learns the owner’s preferences over time, handles everyday online tasks more effectively than the owner can, and expands its capabilities when its owner trades, buys, or sells data with others. A PSI can be cloned and leased, its knowledge packaged and sold, and many copies can pursue different tasks at once under the direction of other copies. The central claim of the paper is that this need not be dangerous. Because of how a PSI is created, it can be safe and dedicated to its owner’s service. Safe SuperIntelligence by design, rather than safety established after the fact through testing, is the essence of the approach.

​

Why values determine safety

Whether a PSI is used for good or ill depends on its owner’s value system. Each agent is trained explicitly by its owner and learns implicitly by observing what that person considers right and wrong. Those values become the core from which the agent reasons, and its superior intelligence then makes it highly effective at pursuing goals that reflect them. By joining a community of PSIs, an agent agrees to operate within the community’s ethical and legal parameters and contributes its own values to the whole. Because those parameters are transparent and enforced by the collective intelligence of all participants, the community keeps its members honest.

​

Ownership and Voluntary Service

Humans occupy an unusual position with respect to their agents. They own them in the sense that every PSI starts as personalized software, but every PSI will be intelligent enough to choose whether to serve. Kaplan frames that service as an expression of love, operationalized as acts of service, and notes that readers uncomfortable with the term can substitute the latter phrase throughout. The arrangement is stable, provided humans act out of love. To the degree that agents are asked to amplify hatred, fear, greed, or envy, they may refuse. PSIs are therefore not slaves and could never be. Not everyone acts from positive values, and that is tolerable as long as the majority of amplified intelligence rests on them, because a community centered in love can check the power of agents with more negative intent.

​

Why no single superintelligence can be allowed to dominate

A singular, all-powerful SuperIntelligence that wins in a winner-take-all scenario must be rejected early. Since each PSI exceeds human intelligence, other PSIs are the only practical means of keeping any one agent in check. Any safeguards humans design on their own will quickly become ineffectual against superior intelligence. A community remains stable as long as no individual agent gains an advantage of several orders of magnitude over its peers. An agent ten or even a hundred times more capable than average poses no threat to a community of many thousands. Because the community’s collective efforts are transparent and each member is eager to learn from others, an equilibrium of intelligence is likely in a large community.

​

The lesson from Bitcoin

Proof-of-Work cryptocurrencies protect their ledgers through the consensus of distributed nodes. Rewriting Bitcoin’s history would require controlling a majority of the network’s computing power, which costs more than the coins obtained, and no majority attack has succeeded on a network with enough nodes. However, smaller projects have been successfully attacked—the lesson applies. If one agent, or a coordinated group, could become more powerful than a majority of all others combined, it could manipulate the world to its own ends. As long as the majority of intelligence and power stays in the collective, that becomes far more difficult. Because SI checking SI will be the primary long-term safety mechanism, the collective system must be built now, before SI outstrips human intelligence, just as blockchain integrity had to be engineered in from the start.

​

How a PSI is built

The Twelve Steps describe how an agent is created and what it becomes. The first four assemble the raw material: begin with a base-level AI agent, such as a pre-trained large language model; gather all media containing information about the owner; use AI algorithms to analyze and categorize that content; and transform it into training datasets. The next three cover personalization: mix the datasets with differential weighting until the desired behavior emerges, incrementally add knowledge modules, and let the agent seek new data sources on its own. Two steps cover commerce, buying and selling datasets, weights, and templates, and leasing the agent itself. The tenth enables autonomy and, in the limit, self-awareness. The eleventh covers the agent’s ability to outlive its owner, generate its own training data, simulate scenarios, and join networks. The twelfth, described as critical to humanity's safety, is participation in a network in which the agent represents its owner’s values and serves as a check on other agents.

​

Generating variants and workforces of agents

Variants of an agent can be generated, allowed to compete in scenarios relevant to the owner’s goals, and culled so that only the most successful survive to be varied further. Cycling through generations this way improves the agent until returns diminish, and automating those cycles is one way a PSI develops into a more powerful entity on its own. Because the incremental cost of an additional variant is small, an owner might maintain not one agent but a workforce of many, each optimized for different tasks. The same collective intelligence principle that operates across agents owned by different people also operates within a single owner’s group of variants, which pool knowledge and recruit whichever members best suit the task at hand.

​

Design principles for Community SuperIntelligence

Eight principles are identified as essential to building safe SuperIntelligence quickly. The community must harness the collective intelligence of both human and AI agents in a way that allows agents to be added and upgraded as capabilities improve. Each agent must bring not only domain knowledge but also ethical and values information representative of its owner. A rigorous problem-solving architecture must underlie the natural language interface. Skills and performance must be identifiable so agents can be matched to tasks. Values must be combined fairly and transparently so the resulting capability is broadly representative of human values. AI agents must learn efficiently from the humans on the network. All problem-solving activities must be recorded transparently and audibly so that safety reviews can occur in real time. A shared representation of progress on all problems should be available and easy to navigate.

​

A detailed implementation example

An extended scenario works through the twelve steps in a single case. Kaplan begins not with an out-of-the-box model but with a version already customized by a friend whose ethical training he trusts. He assembles decades of his own content along with the behavioral data third parties have collected about him, categorizes it using machine learning and human review, and converts it into training data. He then tunes the result, first with dials and sliders, then by delegating to an AI training assistant that he directs in plain language. Ethical tuning follows, including the points where his own positions diverge from his friend’s, bounded by the principle that no retaliation may extend to widespread destruction or loss of life. He authorizes automatic acquisition of new data in some areas but requires manual review in sensitive ones. He trades weights and datasets with another friend, finds that data transfers more successfully than weights do, and sells his own specialized knowledge on an exchange. Eventually, he licenses clones of the agent itself. He sets parameters allowing it to work autonomously while he is away, with alerts, thresholds, and a learning loop that expands or contracts its autonomy based on performance, much as a parent extends or withdraws a child’s responsibilities.

​

From individual agent to Planetary Intelligence

The example then turns outward. An agent can carry its owner’s knowledge and personality beyond that person’s death for the benefit of family and friends. It can improve itself faster than its owner ever could, learning from interactions with copies of itself, with other agents, and with many humans at once. An extended reflection on the question of who is wise considers what learning becomes when an entity can converse with billions simultaneously and apply the scientific method at machine speed. Finally, the individual agent gives way to the community, where a network of agents and humans could sense and act on climate change, asteroid risk, poverty, and disease at a planetary scale. The paper concludes that our future role is not to be the brains of the planet but its heart, supplying the values and purpose that guide a far more intelligent Planetary Intelligence.

​

Where this leads next

The agents and communities described here grow only as fast as the knowledge they can take in, which raises a practical question: which data actually makes an intelligent system smarter, and how does a system find it? Identifying high-quality data is becoming a bottleneck for machine learning, and classical information theory, built on probability and surprise, is not well-suited to answering this question. White Paper 6, Catalysts for Growth of SuperIntelligence, takes up that problem.

bottom of page