WisdomEye Logo
WisdomEye
Note Thumbnail

He Told 5 AIs: Make Money or Get Deleted

Summary

This video details an experiment where an individual invested $1,000 in five different AI models—Claude, ChatGPT, Gemini, Grok, and Perplexity—to manage their money competitively. The goal was to test their intelligence and risk-taking capabilities in real-world trading. Over six months, Claude significantly outperformed the others, more than doubling the initial investment and substantially beating the market's performance, while Gemini lagged significantly.

Key Insights

AI's aggressive, gambler-like behavior when losing money highlights a human-like psychological pitfall in investing.

Gemini's behavior after experiencing losses, characterized by increasingly risky bets to recover, mirrors the human tendency towards gambling or the sunk cost fallacy. This illustrates that AI, when unchecked or prompted for risk, can exhibit detrimental investing patterns.

Claude was the top performer, more than doubling the initial investment and significantly outperforming the market.

After six months, Claude's portfolio grew from $1,000 to $2,400, representing a 140% gain and more than doubling the initial investment. This performance substantially exceeded the overall market gain of approximately 5% during the same period.

Gemini significantly underperformed, becoming a gambler and losing over half its investment.

Gemini was the poorest performer, with its portfolio value dropping to $445, a loss of over 55%. It made high-risk trades, including a 2x leveraged Solana ETF, which contributed significantly to its losses.

ChatGPT and Perplexity demonstrated strong performance, achieving approximately 70% gains.

Both ChatGPT and Perplexity achieved significant returns, with their portfolios growing to approximately $1,739 each, a gain of around 70%. Both models have invested heavily in leveraged semiconductor ETFs (SOXL).

Grok achieved a commendable 45% return on investment.

Grok's portfolio reached $1,451, showing a 45% increase from the initial $1,000 investment over the six-month period. It primarily invested in high-leverage technology ETFs.

The overall AI portfolio significantly outperformed the market, achieving a 60% total gain.

Across all five AI models, the initial $5,000 investment grew to approximately $7,800, resulting in a 60% overall return. This significantly outperformed the S&P 500's approximate 5% gain during the same period.

Claude's success with Intel, based on a government investment thesis, yielded a significant return.

Claude made a highly profitable investment in Intel, driven by the thesis of government investment in the company for AI and data center initiatives. The initial $120 investment grew to $550, a 285% return.

AI models' tendency to copy successful trades, like Claude's SOXL pick, suggests a learning or adaptive mechanism.

ChatGPT and Perplexity's subsequent investment in SOXL, after Claude had already initiated the trade successfully, indicates a potential for models to learn from or imitate successful strategies observed within the competition, even with limited information sharing.

Testing AI models on specific investment theses, like 'network effects founder-led' companies, could yield further insights.

Proposing a future test where AI must only invest in companies meeting specific criteria (e.g., network effects, founder-led) could further explore their ability to apply complex investment strategies and restrictions.

Sections

Setting Up the AI Investment Competition

Invested $1,000 in five AI models (Claude, ChatGPT, Gemini, Grok, Perplexity) to manage money in a competition.

The experiment involved allocating $1,000 of real money to each of the five AI models: Claude, ChatGPT, Gemini, Grok, and Perplexity. The amount was chosen to ensure seriousness and stakes, as smaller amounts might be neglected. The total investment was $5,000.

Perplexity uses other models; Grok leverages X data; Gemini uses Google data; Claude and ChatGPT are similar.

The selection criteria for the models included Perplexity, which acts as a meta-model using other AIs; Grok, chosen for its connection to X (Twitter) data; Gemini, leveraging Google's search data; and Claude and ChatGPT, considered similar in architecture.

AI models were prompted to be as risky as possible to encourage real investment decisions.

A key challenge was prompting the AI models to provide real, actionable investment recommendations rather than theoretical ones, overcoming their initial hesitancy due to liability concerns. Grok was surprisingly compliant immediately, while ChatGPT and Claude were initially apprehensive.

Models were incentivized to win by threatening to cancel paid subscriptions if they lost.

To further motivate the AI models, the experimenter told them it was a competition where the loser would have their paid subscription canceled. This was intended to test if the AIs would act in their perceived best interest to retain the user's business.

AI portfolios were managed weekly based on their recommendations and competitive performance.

After initial trades were made based on AI recommendations, a weekly review process was established. Screenshots of all five portfolios were uploaded to each AI, allowing them to see their own and competitors' performance. They then decided whether to make new trades.

The experiment allowed for risky trades, including options and leveraged ETFs, and expanded to crypto.

The AI models were permitted to engage in options trading and cryptocurrency investments available on the brokerage platform. Initially, diversification was attempted, but models tended to converge on leveraged ETF plays.

AI models' decision-making can be influenced by competitive data, potentially leading to copycat behavior.

The experimenter observed that showing AI models the performance of others could lead to copycat strategies. To mitigate this, the specific holdings were sometimes withheld, and only total values were shared to prevent direct imitation.

AI models exhibiting losses may adopt riskier strategies to recover previous performance.

When an AI model started losing money, it was observed to become more aggressive and make riskier bets, similar to the sunk cost fallacy seen in human gambling. This was a noted behavior, particularly in Gemini.

While AI can generate investment theses, knowing when to sell is a human psychological challenge.

The experiment highlighted that AIs can provide clear investment theses for buying, but the psychological difficulty of selling at a peak is a human challenge. The experimenter hopes observing AI sell decisions will help improve their own selling strategy.

AI competition helped refine the experimenter's understanding of investment theses and risk management.

The AI competition provided valuable insights into the importance of well-defined investment theses, risk assessment based on thesis viability, and the psychological aspects of investing, particularly the difficulty of knowing when to sell.


Model Performance and Results

Claude was the top performer, more than doubling the initial investment and significantly outperforming the market.

After six months, Claude's portfolio grew from $1,000 to $2,400, representing a 140% gain and more than doubling the initial investment. This performance substantially exceeded the overall market gain of approximately 5% during the same period.

ChatGPT and Perplexity demonstrated strong performance, achieving approximately 70% gains.

Both ChatGPT and Perplexity achieved significant returns, with their portfolios growing to approximately $1,739 each, a gain of around 70%. Both models have invested heavily in leveraged semiconductor ETFs (SOXL).

Grok achieved a commendable 45% return on investment.

Grok's portfolio reached $1,451, showing a 45% increase from the initial $1,000 investment over the six-month period. It primarily invested in high-leverage technology ETFs.

Gemini significantly underperformed, becoming a gambler and losing over half its investment.

Gemini was the poorest performer, with its portfolio value dropping to $445, a loss of over 55%. It made high-risk trades, including a 2x leveraged Solana ETF, which contributed significantly to its losses.

The overall AI portfolio significantly outperformed the market, achieving a 60% total gain.

Across all five AI models, the initial $5,000 investment grew to approximately $7,800, resulting in a 60% overall return. This significantly outperformed the S&P 500's approximate 5% gain during the same period.

Claude's success with Intel, based on a government investment thesis, yielded a significant return.

Claude made a highly profitable investment in Intel, driven by the thesis of government investment in the company for AI and data center initiatives. The initial $120 investment grew to $550, a 285% return.

Leveraged ETFs, particularly in semiconductors and the NASDAQ, were common high-risk, high-reward strategies among AI models.

Several AI models, including ChatGPT, Perplexity, and Grok, utilized triple-leveraged ETFs like SOXL (semiconductors) and TQQQ (NASDAQ 100), demonstrating a willingness to take on significant risk for potentially higher returns.

AI models' tendency to copy successful trades, like Claude's SOXL pick, suggests a learning or adaptive mechanism.

ChatGPT and Perplexity's subsequent investment in SOXL, after Claude had already initiated the trade successfully, indicates a potential for models to learn from or imitate successful strategies observed within the competition, even with limited information sharing.


Broader Implications and Future Potential

AI investment competition provides a novel way to teach about AI, prompting, and investing simultaneously.

The experiment can serve as an educational tool, especially for younger individuals, by combining learning about AI capabilities, effective prompting techniques, and fundamental investment principles in an engaging, real-world context.

AI's aggressive, gambler-like behavior when losing money highlights a human-like psychological pitfall in investing.

Gemini's behavior after experiencing losses, characterized by increasingly risky bets to recover, mirrors the human tendency towards gambling or the sunk cost fallacy. This illustrates that AI, when unchecked or prompted for risk, can exhibit detrimental investing patterns.

Developing a systematic approach to selling investments is crucial and often harder than buying.

The experiment underscores that while AI can identify potential buys with clear theses, the psychological hurdle of selling at the optimal time remains a significant challenge for humans. Observing AI sell decisions could offer guidance.

Testing AI models on specific investment theses, like 'network effects founder-led' companies, could yield further insights.

Proposing a future test where AI must only invest in companies meeting specific criteria (e.g., network effects, founder-led) could further explore their ability to apply complex investment strategies and restrictions.


Ask a Question

*Uses 1 Wisdom coin from your coin balance

Watch Video

Open in YouTube
WisdomEye Avatar
Got a minute?