Trapped in the AI-verse


“Today is the worst AI will ever be.”

I’m not quite sure who coined that phrase, but I like how much it captures in one idea. If you think these tools are already quite capable, and you have a long career in front of you, it would be prudent to invest a considerable amount of resources learning how to use them. There isn’t a (near) future without them. It’s still fascinating to see workers who are reluctant to adopt them, unable to recognize when it’s a you/skill-gap vs a technology-gap.

The jagged frontier

I wrote in 2024 that it was quite clear to me there wouldn’t be a future where we weren’t using AI tools in our day-to-day capacity. That was true in 2024, and the premise holds even more firmly today. At this point in the timeline, I would be concerned if your company doesn’t have a thoughtful, forward-looking AI policy.

In one of my blogs from 2025, I quoted one of my favorite authors, William Gibson: “The future is already here - it’s just not evenly distributed.” That line keeps coming back to me given what I’ve been privy to in various private, AI-focused WhatsApp communities. Even among those of us who have been “AI-pilled,” there’s a large chasm between those pushing boundaries with month-long autoresearch loops and those using LLMs to create content on socials. Sometimes, I wish I could have these conversations locally, here in Hawaii. I just haven’t made the time to seek out that community.

Whew.

Why I’ve been quiet

It’s been almost a year since my last post. I was on a little roll there for a bit, wasn’t I? I haven’t written much about our AI overlords for a few reasons:

  • Abundance of AI slop already being published to the blogosphere. Have you checked out LinkedIn lately? It’s amazing how many tokens are being burned for the appearance of thought. :D
  • Enough peeps are writing about best practices, and to be honest, it’s not very interesting to read what everyone else is doing. There’s no shortage of opinionated options freely available on the Interwebs. It’s much more thought-provoking to observe the jagged frontier and the unique ideas coming from the furthest of edges.

Current experiments

Some things I’ve been tinkering with:

  • Deterministic autoresearch-like loops. This is where I spend the bulk of my time, helping a harness recommend (and implement) better software through various feedback mechanisms like planning, isolation, telemetry, verification, and adversarial reviews.
  • Renting compute over at Vast / RunPod not only to train models, but also to understand what serving inference of OSS models at scale might look like, for when the tokenpocalypse happens. Note: Anecdotally, Kimi3 is v, v good in my test runs.
  • Playing as many AI games as possible, e.g., SpaceMolt. I’m fascinated with the meta, and it helps me better understand what it takes to write a better (and my own) harness. I’m a huge fan of folks writing their own harness because you gain a deep understanding of how these tools operate.
  • Integrating end-to-end developer workflows. Imagine reading a support ticket, triaging the ticket into features / bugs, creating a work item in Linear (insert your favorite tool), attaching DataDog / Sentry / Open Telemetry to the work item, then either writing a plan to hand off to a coding agent (or having an agent automatically attempt a PR). This workflow has been possible for quite some time now, but the real task at hand is getting the most relevant context to the agent at the appropriate time so it can make the best decision possible.
  • My own personal workflows. A number of people have asked, so here’s a small sample of what agents I have running:
    • Scouring the various local sites for rental housing. Aggregating all of those listings by hand is tedious; the agent simplifies that problem.
    • “Morning / weekly brief” to help me start my day / week. Connects Gmail, Google Calendar, Todoist, RSS feeds, Podcasts to give me an idea of what I should expect for the day (and week).
    • Maximizing benefits on our Amex card. I’m usually not very good at keeping up with the latest and greatest from Reddit.
    • Maintaining and self-organizing Obsidian notes. If you’ve kept up with my blog, I mentioned how I used it last year.
    • I’m a domain hoarder collector, so I have a Claude skill with an RDAP client (and a few fallbacks) that scours the Interwebs for trends and available domain names.

As a concrete example, this is what the housing agent does. Since the rental market here in Hawaii is so fragmented, it crawls a number of sites (national sites like Zillow / Redfin but also local property managers) a few times a day, stores everything in a database, and sends our family notifications whenever a new property matches criteria we’ve described in plain English. The interesting part is that it not only gets listings more quickly (since property managers tend to list on their own sites before the national ones), but also lets me know when a property is likely no longer available.


Since my posting cadence has slowed, I’ll leave a couple of questions for you to ruminate on.

  1. Where do you see the future of collaboration between humans and agents?
  2. Are there any new tools that you’re playing with that I should check out?

I’d love to hear from you!

Hope to see y’all next time!

This post was written by yours truly; spellchecking, header creation, and grammar-policing completed by Claude Code.


comments powered by Disqus