execution stopped being the bottleneck this summer — judgment did
How I actually use agents
Well, summer 2026 is coming to a close, and I am so grateful I’ve been able to travel to see good friends and make a lot of new ones. I’ve been to Foo Camp in Berkeley, the 2389 Hack Party in Montreal, YC Startup School in SF, Memphis for family, and now I’m back in New York (home).

Every time I leave great conversations with a high density of people who are thinking about the same things I am, and then step out of that environment, I still get culture shock. It’s interesting how many people talk about AI versus how much they actually integrate it into their daily lives to handle real workflows.
There are heavy users of different agent harnesses like OpenClaw, Hermes, Pi, etc. Now we’re also seeing a wave of organizational and company “brain” harnesses. Jack Dorsey has Buzz. YC built their own Request for Startups harness, QM. Cursor and xAI came out with Grok bot. Everyone is slowly, but still very quickly, arriving at some version of humans and agents working together.
Claude Tag really accelerated this for everyone this summer.
It made human–agent multiplayer feel less theoretical. The adoption is that it sits where a lot of organizations already are which is Slack. Other products quickly brought similar ideas into more “normal” work environments.
Maybe that was the moment.
Maybe that’s when people finally started seeing what AI could actually be used for.
But I keep getting the same question:
How are people actually using personal agents in their daily lives?
Every time someone asks me that, I’m still a little surprised. I have to keep reminding myself that the people I’m around most of the time live in this world, and not everyone else does. To be outside of the AI bubble is to not see any of this as obvious.
I’m around people who make things happen. We have an idea for something and then we (with agents) hack a version of it, and it just works for our lives.
Something can feel obvious because everyone around you already has built some weird version of it.
A lot of things that feel simple, hackable, or almost comically meta inside this bubble are completely foreign outside of it.
We’ve moved so fast that when someone asks, “What are others doing with their agents?” I’ve started bringing the question back to myself:
How am I best using agents?
I’m using the Claude Code harness for agentic coding.
The productivity increase is obvious. I can build more, test more ideas, and move faster. But more productivity also creates more abundance.
More code.
More ideas.
More things I could be doing.
I have come to the conclusion that execution stops being the bottleneck and judgment does.
So I’ve started to write more, read more, say my thoughts out loud, and become more introspective.
What do I believe?
Has my opinion changed?
This summer I’ve been thinking about “taste” and what taste actually means.
The best way to give something your taste is to articulate it.
For me, that means writing and talking into my phone.
Which is funny, because since ChatGPT arrived with the potential of supermaxxing productivity, writing started to feel optional. Why write something yourself when you have a “thing” that can write for you?
But how can you even be productive and tell an agent what to do when you don’t know what you want to do with agents? The intent is missing.
So now I keep going back to a much simpler question:
What is literally on my mind right now, and how do I portray that to my agent?
That question pulled me deeper into context management and memory. I did some deep dives and wrote about my findings last fall.
But I want more, and I want to do better. As I do more and talk to others, I dump all of my meeting notes, voice notes, and random thoughts into my agents.
For feedback, it’s great. It surfaces converging themes. But then I question myself again: What are my own views? What are my beliefs? Where is my conviction?
How can my agents keep up with what I think, how I think, and why?
So I’m putting it to the test with the most powerful models today.
YC Startup School experiment
After YC SUS, I made a judgment graph with the Bestmate harness. I put my notes and the speaker notes into it and used the harness to identify all of the converging themes. I could go into each point afterward, run a Q&A with each “speaker,” and trace where the idea came from.

I then wanted more of a multiplayer world of seeing what other attendees thought, so I sent it to my new friend Fiona and they added theirs as well.


I keep going back to this every time I have a new batch of speaker notes, meetings, or research. I want to see how I can apply these ideas to my life right now to better myself overall.
Where did I agree and disagree?
What changed my mind, and why?
Which ideas keep showing up in my decisions?
Where am I contradicting myself?
What do I believe more strongly now than I did six months ago?
Isn’t that the whole point?
We’re building self-learning loops and verification systems. Is it because we want AI to be more human? To think like us, switch personas and tone, trace decisions, surface intentions and patterns, give us different viewpoints, and do actual work?
That’s the more meta layer, and more and more startups are breaking it down, attacking it from different angles.
Memory is raw material for judgment. A memory system remembers: “Kaya went to a talk where someone said X.”
A judgment system should eventually be able to say:
“Kaya has now heard versions of X from five different people. She tends to agree with the premise but consistently rejects one assumption underneath it. Her position appears to have shifted after these three conversations.”
That is more interesting to me.
And it changes what something as simple as a daily brief can become. Most weekly AI summaries tell you about the meetings you had, the tasks you completed, the things you talked about which are all useful don’t get me wrong.
But what if there were a weekly brief that said:
“You encountered the same argument about AI teammates in four different conversations this week. You agreed with it twice and challenged it twice. The distinction you seem to care about is whether an agent has access to information or actually understands the judgment behind that information.”
Now that is definitely more interesting. The agent is not just remembering my week; it’s watching my thinking evolve.
I made a Maslow’s hierarchy of needs for agents. At the bottom are the basics; access and integration. Plugging into the places work actually happens; Slack, Teams, etc. Above that is memory and context: remembering conversations, preferences, prior decisions, and the running state of my world. Then comes the execution and delegation: taking action, completing workflow, handing work back and forth between humans and agents.
But at the top of the pyramid is where it is really interesting for me: reflection, traceability, and finally judgment and taste.

That’s the layer where an agent doesn’t just recall what happened but understands what matters to me (beliefs, priorities, evolving opinions, and decision patterns) and then can reflect those back, show me contradictions, and show how my thinking is changing over time.
Agent Interfaces
Maybe it’s all the same ideas, but the interface truly matters.
Because the interface determines what kind of context an agent can actually collect from you.
If I only interact with an agent through a chat box, it mostly knows what I consciously decide to tell it.
If it sits inside my meetings, it knows what I heard, what I said, what I questioned, and what I pushed back on.
If it sits inside my work, it can see what I prioritize, what I delegate, what I ignore, and what I keep coming back to.
And if I give it an interface for tracing judgment, now it can start understanding the why behind those behaviors and it’s even a better way for me to traverse it.
That is why I think where agents live matters so much. Slack is powerful because work already happens there. Granola is useful because conversations already happen there. A judgment graph works because it gives me a way to explicitly react to ideas instead of just storing them.
The interface is not just how we talk to the agent. It determines what version of ourselves the agent gets to know.
Maybe this is also part of why people still ask how to use AI in their daily lives. A blank chatbot asks you to already know what you want. Most people don’t and it’s even harder to convey that visually. The more interesting interfaces sit inside places where behavior is already happening: meetings, Slack, work notes, web, and let the agent learn from that behavior instead of waiting for you to manually explain yourself.
But the important part is that I don’t want the agent silently deciding what I believe. I want it to form a hypothesis about my judgment, surface that back to me, and let me correct it.
Maybe it notices that I keep pushing back on the same idea across different conversations. It should be able to say, “I think this is what you believe and this is why,” and then give me the opportunity to say yes, no, or not quite.
That loop matters.
Observe. Infer. Reflect. Correct.
I’m still in the loop.
From this point of view of proactive context, I have more confidence using my agents in outward-facing contexts for work because they now have my judgment embedded. They can pick up on the same contradictions I battle in my day to day decisions.
How we see this today
You’ve probably seen many versions of this: weekly wraps, daily briefs, cron jobs, POV articles, etc.
What matters is that it actually works once the underlying judgment system is in place.
I have my agents build weekly wraps based on my judgment to see how it evolves over time. Because at the end of the day, it’s fine-tuned on me. I’d like to see a future of personalized LLMs that can be outsourced to others and used however we want; carrying our judgment with them.
From there, you can move from personal context to “work mode” while still conveying your judgment throughout. That’s where I believe AI teammates, company harnesses, and the true multiplayer world of humans and agents really come into play, but it has to start with the self.
You can play with iterations of this and plug it into whatever harness you want to really live in the agentic world:
Judgment graph.
Everything starts with getting raw thinking in with as little friction as possible.
My meetings flow in automatically through Granola. Everything else is one move. I can ingest a link, drop in a file, or drag a document into Bestmate and decide where it belongs.
Then I build the graph.

The first pass clusters everything into themes and claims. But the part that turns it into a judgment graph instead of a notes graph is the interview.
Bestmate asks me what I actually think about the claims it found.
My answers get written back into the knowledge base, and the graph rebuilds around both what I consumed and where I landed on it.
So what comes out is not just a map of information.
It is a map of my relationship to that information.

A nightly process reclusters everything so what I said, what I decided, and what the system inferred about me do not slowly drift into separate versions of the same thought.
Tracing your judgment
A judgment without provenance is just another model guess. Every belief in the graph should be able to answer one question:
Where did this come from?
When I trace a judgment, I can see the actual evidence behind it. Which meeting, which email, what date. The exact text in my knowledge base. I can open the full source instead of only seeing a generated snippet.

The correction loop I talked about earlier becomes something concrete here. If the system gets something wrong, I should be able to correct it directly.
Incorrect → Trace this → Where else is this mentioned → Forget this.
And a correction should do more than hide a sentence. It should invalidate the belief, flag where it came from, and stop the system from quietly rebuilding the same wrong conclusion later.
That is also what makes the idea of watching my thinking evolve real rather than just a hypothetical.
If I paste in something I am thinking about, Bestmate can give me back my own take based on the positions I have accumulated over time.
And if what I am saying now conflicts with something I decided months ago, I want it to tell me.
Most AI agrees with you.
Bestmate tells you when you are contradicting yourself.

Then every week, Wrapped shows me the diff of my own mind.
What changed.
What I reinforced.
What I contradicted.
What I no longer believe.
What replaced it.
The weekly brief I used to describe hypothetically is now something I can actually run against my own judgment.
Giving other people your judgment.
Once judgment can be captured and traced, the next question is whether it can travel.
I think about that at three different levels.
The tightest level is agents.
I can give Claude Code, OpenClaw, Hermes, or another agent access to my judgment before it acts. When one of them reaches an “it depends” moment, instead of guessing, it can send the question back to me. I answer once. That answer becomes durable judgment that every connected agent can inherit.
Answer once, teach all of them. The middle level is people.
Someone should be able to interact with a version of my judgment without getting access to my entire private context.
I can invite someone into a workspace or share my twin — I share twins where I work (Slack, Telegram, iMessage, WhatsApp). They can ask it questions, see how I would approach something, and escalate back to me when the system is unsure. My answer then becomes part of the judgment layer. But the boundaries still matter and the fences should hold by construction, not by policy.

Then there is the loosest level.
The world.
I can publish a judgment graph and let anyone interrogate the thinking directly. At that point, what is being shared is no longer just a static essay or document. It is a living body of reasoning.
Not a clone. Not an AI wearing my name. Just my reasoning, made legible enough to hand to someone else. First to my agents. Then to the people I work with. And, sooner than I expected, to their agents too.
Bestmate is in beta. If you want to feel what the homepage of the agentic computer is like on top of your own world you can book a demo here.
© 2026 Forever 22 LLC. All rights reserved. The tech insights and personal experiences published here represent the operational findings of Forever 22 LLC and the author. Sharing via direct links is highly encouraged, but no text or code may be reproduced or used without explicit written consent.