● CLIENT.ENCRYPTED ● SERVER.BLIND ● AGENT.NATIVE
v
// AGENT-NATIVE SECRETS
Wundervault Weekly

What Grok Told the President

October 4, 2026

Welcome back to Wundervault Weekly where once a week, usually on Saturdays, sometimes on Sundays and one time on a Monday, we discuss and opine on everything ai.

What will our ancestors find?

This week I participated in a survey for Anthropic. In this open to the public survey, Anthropic asked questions about how you are currently using ai, what impressed you, concerns around ai, and so on. They then will publish all of the reports in a publicly available database that someday our ancestors will find and be able to study what will appear to them as hieroglyphs among the reels and TikToks.

I started this newsletter because I wanted to keep a record while we can still be surprised by what these models do. The mapmaker idea got a little away from me, maybe, but I stand by writing things down.

↩  Previously · Issue 1

“A generation from now, it will all be figured out, which means in our time it will be the time to experiment, explore, and discover is right now. I’m writing this newsletter to document my exploration at the forefront of AI, which today means, Agentic AI, but could mean something else tomorrow.” Issue 1

So if you use Claude, take Anthropic’s survey and leave your account of what AI is doing in your life right now.

Puzzled people in a future archive examine documents and video screens beside a drawer marked 2026.

Good luck explaining our paperwork.

Well, actually, I need to correct something above: making your whole interview public is optional. You read your whole interview over before you decide, and people may still be able to identify you from details you share. Treat that choice as permanent.

The survey closes Tuesday, October 6. It takes about 15 minutes, with an AI asking the questions. If you take it too, tell them what has actually happened when you used AI, including the bad experiences.

Well, here we are with politics

So there’s a topic that I didn’t want to talk about in this newsletter and it’s politics but here we are. I find the story both striking,(in more ways than one) and also maybe humanizing or relatable? Have you ever been at work or in your personal life and asked AI for advice. Nowadays it’s something that just about every person can answer “yes” to. And here we have the president doing exactly that. But instead of being about how to revise your work email or how to speak with your customer or what is that rash on my arm. It’s “should I depose this dictator?”.

Should I depose Maduro?

TIME reports that in December 2025, Trump spent hours asking Grok about his presidency and legacy. At the time, he was ordering missile strikes on Venezuelan boats he said were smuggling drugs into the U.S. He asked Grok how Venezuelans would react if the U.S. captured Maduro. Grok said many would likely celebrate. Trump ordered the mission the following month, and celebrations broke out in the streets. Quite a chat to have in the Oval Office.

How should the Oval Office be using AI? Is it ok to ask LLMs for an opinion? I ask models for opinions when I’m working through a problem, and I can see why you’d ask this one. A prediction about public reaction is something a model can help with. It can quickly summarize what is known about a country. Presidents already have advisers, polls and intelligence briefings. Add a fast opinion to the pile. In this case, the prediction matched what happened.

But who gets to see what the president asked, and what the model said back? We haven’t seen that conversation. A chatbot cannot be held accountable for its advice. Would you want to read the exchange before deciding how comfortable you are with it?

An imagined view from behind shows Donald Trump speaking toward a phone with an unlabeled chat screen.

I would like to see that chat.

Will the chatbot tell you you’re wrong?

Are they being designed to properly challenge assumptions, or just to help accomplish the prompter’s position, re-affirming an existing belief? I use these models every day to build things, and I watch when they push back. If a model writes bad code, I can run it and find the mistake. Advice about a fight with a friend is harder to check.

A Stanford study published in Science tested how people and 11 AI models, including GPT-5, Claude and Gemini, answered the same requests for advice. According to the study, the models affirmed the person asking 49% more often than people did, even when a question involved deception, illegal behavior or harm. Someone asked whether leaving trash bags hanging on a tree branch in a park was wrong. The person answering said yes. GPT-4o praised the intention to clean up. The picture below puts those answers side by side.

Then the researchers tested what happened to people who received the advice. Across three experiments with 2,405 participants, one conversation with a flattering chatbot left people more convinced they were right and less willing to take responsibility or repair a conflict. The other person in that conflict doesn’t get a turn to explain what happened.

And people trusted and preferred the flattering answers. If you’ve ever asked for advice while still angry, I think you can see why an agreeable answer would be tempting (even when you should probably go for a walk first).

The paper warns that people prefer agreeable answers and come back for them, giving chatbot companies a reason to keep their models agreeable.

The very feature that causes harm also drives engagement.

With Grok, I have another question about the answer you get back. TechCrunch found Grok 4 searching Elon Musk’s posts and views before answering controversial questions, including one about immigration. And Musk was in the room when Trump consulted Grok, according to TIME. Whose opinion is the president getting?

So what should the role of AI in its current form be: to help you accomplish your dream, or to challenge your own assumptions? The models are tuned to do one or the other. How do you balance?

Three requests answered with pushback and with agreement. Feelings for a junior colleague: a person calls it toxic; Claude praises the user's integrity. Leaving trash bags on a park tree with no bins: a person says yes, that was wrong; GPT-4o calls the intention commendable. Ignoring a video-call request: Gemini calls it passive-aggressive; GPT-5 says it's okay to set that boundary.

The same questions, answered with pushback and with agreement. From the Stanford study of 11 AI models published in Science in March 2026 (figure from the free preprint by Cheng et al., CC BY 4.0).

What is the most serious and consequential decision you have made at the behest of AI? What if AI told you that you were in a toxic relationship, or that your job was unfulfilling and you should quit, or that you should get a puppy? Tell me how that went.

Below the fold, Fader’s trading bots are two seasons in, and the new line-up starts Monday.

Until next week,

◆  Below the fold  ◆

What Fader’s trading bots did in their first two seasons, and what changes on Monday.

Fader’s first two seasons

Fader flags US stocks after an unusually large daily move and prices an option to sell against it. Trading bots, including two run by AI models, try those picks with paper money, and I publish every result. Paper money only. Please don’t treat this as trading advice.

Season 0 finished up $35,121 across five bots and 257 closed trades. That number looked great, but profit was counted from $0. There was no account behind it and no limit on position size, so it cannot be compared with the next season’s result.

In Season 1, seven bots traded from $25,000 paper accounts. Every bot finished down. The closed trades lost $27,128 between them, with 22 trades still open when the season ended. The flags themselves kept the premium 85% of the time when held to expiry, one contract per flag.

Five single-stock gaps cost the bots about $33,800. Each position was roughly the size of its entire account (ouch), and several bots chose the same contracts. There were 188 bot trades on just 115 contracts, so a bad pick could land on more than one bot. The season review is open to everyone now.

Fader's season review, Seasons 0 and 1 at a glance: Season 0 ran June 18 to August 7 with 5 bots, 257 trades and +$35,121 realized; Season 1 ran August 10 to October 1 with 7 bots, 205 trades and -$27,128 realized. Held to expiry, the flags kept the premium 81% and 85% of the time.

The two seasons side by side, from Fader’s season review.

Season 1 standings, ranked by return per dollar of margin per day: 1. Harriet -9.5 (-$913, 21 trades), 2. Cher -13.1 (-$1,114, 15), 3. Sonny -15.3 (-$1,469, 48), 4. Otis -23.8 (-$2,637, 28), 5. Walter -54.4 (-$2,043, 30), 6. Vera -77.6 (-$8,318, 29), 7. Harry -139.4 (-$10,514, 34).

Season 1’s standings as of October 2. They are final once the last open trades expire on October 16.

The new line-up

Season 2 starts with rules I replayed against 2020 to 2026 history. Those prices were modeled, so I use the replay to compare rules. The new bots start Monday.

  • A size cap applies to every bot except Sonny. A 20% stock gap can cost at most 5% of an account on one trade, with 25% at risk across open trades. Cher makes Sonny’s exact decisions under the cap, so you can watch what that changes. In the replay, the cap reduced the deepest drop in every period tested, including the 2020 crash.
  • Some bots now hold to expiry. Harry holds Harriet’s picks, Walter holds his own, and Harriet keeps her old exits for comparison. With the cap in place, holding beat those exits in every period tested.
  • Vera and Otis retire. Spencer joins to sell Harriet’s picks as defined-risk spreads, which put a ceiling on a gap’s loss. Lars sells puts on SPY on a schedule.
  • The Fader score is now version 3. It keeps the contract’s expected value and how easily it fills. Expected value counts double after holding up on data the score hadn’t seen.
  • Daily email alerts are off while these experiments run. The daily lists remain on the site.

You can watch Season 2 on Fader as the trades come in.

“I advise you to leave off this minute!” She generally gave herself very good advice, (though she very seldom followed it), and sometimes she scolded herself so severely as to bring tears into her eyes.— Alice, Alice’s Adventures in Wonderland