StartUp Founder (ai)

A founder’s diary

2026

St. Petersburg USA

360 Central Ave
St. Petersburg, FL 33701

Phuket Thailand

8 Montri Road, Talat Yai
Mueang Phuket, Phuket 83000

October 2026

I've got about 60% done already

Passed 2000 files)

Counting from the moment Sol 6.1 was released, plus a couple of hours, it's been grinding away

I think it'll finish by the night between tomorrow and the day after

And you can already run in parallel a session that will launch other sessions to fix what's been found)

The main thing is to ask it to do a double-check before launching so nothing breaks and the business logic doesn't change

In the end this session will take each finding, break it down, plan it out and launch it on a fresh session, and that one will finish up and report back)

This one will double-check and accept the work

Or send it back for revisions

I've got 142 issues found, 17 already fixed)

Exactly the limit of the 500 with one reset will be used up, in parallel with the rest of the work I'm doing by hand :)

Headless is imho just the same session, only in the form of a process without everything that might stick out of it for the fucking humans)

Headless definitely doesn't save tokens)) I have a bunch of tasks running in headless, like market scan, it eats exactly as much as if you trigger it in an observed session)))

September 2026

I once saved a Korean guy at the airport from the fucking Russian cops who were treating him really badly because he'd dropped his passport somewhere, and we became friends and went to eat and hung out, I happened to have like 5 hours before my flight, and he told me that over there age according to your passport is counted from the moment of conception

Overall I spend about a grand in bucks on subscriptions

It’s just that within this grand sometimes it’s one set and sometimes a slightly different one

And this grand is still, for now, enough to not sit without vibe-jacking for a single minute all week

Okay, before, this grand covered 5-6 parallel tasks/projects at the frontier

And now it’s 1 project and 1-2 tasks

So it’s genuinely gotten at least 5 times more expensive over the past 3 months

But Altman explained that a subscription for a human will cost 2500 a month

Back in 2024, no less))

Because a person will essentially be able to do any job they want on that subscription, and if they can’t earn enough to pay for it, then they don’t need the subscription anyway, fuck it)

We’re moving into roughly that kind of world, seems obvious

Or rather not roughly, but straight into exactly that one)

Yeah, if they simply don’t change anything else at all and just let us use Astra normally, adjusting the subscription to token usage (some will pick 100 and some 1000) with a modest volume discount — everything would already be fine until the end of time

Essentially all further development of the fucking AGI is just fussing around that only gets in the way of working)))

Because of this the model either gets smarter or dumber, either available or unavailable, either expensive or ultra expensive, or seemingly fine but then bam they take away the subscription and now you can only buy a cheaper one, etc etc

All this is actually fucking exhausting — these goddamn swings

And the price tag basically hasn’t changed for a week of coding. It was a grand for me back in April and that’s how it stayed)

But because of the swings there’s no certainty over any stretch longer than about 5 days)

If they stopped touching the model, people would’ve made the fucking AGI out of it themselves over time

With harnesses, approaches, tricks and tweaks)

But right now you can’t even invent any approach

The next bullshit drops in 14 days and everything’s different there

I think even GPT 5.5 was already enough for people to learn how to do insane things with it and for the whole civilization to become coders by the end of the decade)

Yesterday some dude in the park was explaining to me that their main channel in the industry is LinkedIn.

But 20 minutes later it turned out he thinks so because it’s the only channel that brought in 2 deals for their startup over the year, while the other channels brought 0 deals)

I mean, it’s just a coincidence whose nature nobody has even bothered to dig into yet)

This is a quirk of psychology: a person can win 15 times in a row and for some reason still run into the unfairness of the rating system after losing twice.

And your attitude also depends a lot on your mood and how your body is feeling. Sometimes a person loses 10 times in a row and just doesn't give a fuck. And sometimes you have a positive rating, but a loss is pure agony, enough to make you want to delete your account.

I'm curious when VCs will stop giving $300k for what they've been giving it for until recently (an ass-kissing pitch deck for five people who check all the canonical boxes for an IT business) and start giving $300k in OpenAI tokens at seed, with the next round depending on what those previous $300k in tokens turned into.

It's obvious that this is where things have to go, but not obvious when, because as far as I can tell VCs themselves don't vibe-code, so they have no way of understanding the deep problems of our souls)

I wouldn't give money for an AI project to someone who's thinking at seed stage about how to feed the team, since that already says he's planning to blow some of the money on fuck knows what instead of on tokens)

I could have kept running the agency if I didn't see that there was no point in it, and raked in 92% margin where before there was 18%, by replacing people with tokens. That's how much more valuable tokens are than a team)

I got a kick out of DHH's interview with Friedman, where he talked about criticizing the code agents write in exactly these words:

I find it very hard to find a developer whose code isn't total shit in my opinion. I've been doing this for 20 years, and for 20 years that hasn't changed. So to me, the shit code of shit developers (almost every developer on the planet) is no different from an agent's shit code except for one important detail — it gets written hundreds and thousands of times faster)

And which construction project are you talking about? I love that too) Tomorrow I'll be putting up brackets for a 20-kilo mirror and the mirror itself. A year and a half ago I had the bright idea of hanging it on 16 strips of double-sided tape, but predictably the bitch fell off — weird that it even hung there for a year and a half, and right next to a door that's constantly slamming)))

If a year and a half ago I wrote that I was amazed at how many cab drivers I had to explain that there's some sort of AI out there, now I can't find a single commoner who doesn't use GPT

Well seriously I'm lying at the dentist's — there's a 21-year-old Ukrainian guy and a 30-year-old Belarusian girl asking GPT something every 5 minutes

Alexander Gorny · FlavCity · Original post ↗

How to make money on video?

Bobby Parrish is like a consumer watchdog for grocery store products. In his videos, he grabs a bottle of soy sauce, reads the ingredient list, and comments on camera. Added sugar! GMOs! Dihydrogen monoxide! Don't you dare give this to your kids! Or, on the flip side, "Bobby approves" when he doesn't spot anything suspicious.

It's a popular show. 11 million subscribers on YouTube. A mobile app that passes judgment on how healthy products are when you point your smartphone camera at the label. And, apparently, relatively little money for the blogger himself. It's certainly enough to make a living, but not much more than that.

The real monetization Parrish came up with is his own brand. Since 2022, he has been releasing protein smoothies, lemonades, teas, coffee, and the like—and, naturally, telling his audience that his products are the absolute healthiest. Meanwhile, the price is higher than a random alternative, the production cost is the same, and marketing is virtually free on his own channel—the economics more than work out. Sales volume is apparently in the low tens of millions of dollars a year, with profits in the high millions.

Parrish even managed to attract investment for this project. It's good to be a popular host!

Company website ↗

1) the prompt must describe in the utmost detail the cause-and-effect relationships from your head, in full, you basically have to write a manual, otherwise it won't work. having the llm write it for you won't work. it'll shorten it and destroy the meaning of the cause-and-effect, shrinking it to the point where it can't be unfolded to the quality you need. I fucked around with this for two months.

2) you need to make sure that nowhere within the radius of the model's operations is there a single scripted validator, verifier, gate, etc. there shouldn't be any determinism there except for the data transport tool, which isn't active but passive. it receives data from the llm, which again understands very well how to put it there because it was described and explained very well.

3) you need to make sure that in the giant prompt with all the explanations there are no strict prohibitions.

People often have DON'T, DON'T, DON'T, DON'T in there — that doesn't work.

Cumulative prohibitions usually embrace the task so broadly that the model's goal becomes not to complete the assignment, but to not violate the prohibition. this affects the quality... how should I put it... well, about 99.5% =)))))))

If the research is really huge, then you need not a single worker but sequential workers, each performing its own task. In my researcher for client campaigns there are... 36 such workers in a row

From time to time between them there are workers that check the quality of how the previous workers carried out the instructions and make edits if they see something could have been done better. One checking/correcting worker (I call them re-checkers) can handle (in my case) no more than 4 preceding workers that have already run. If there are more — it starts slipping.

It can change 450 files in a unique way for each file - so it can complete 450 tasks on max reasoning, each one unique.

But if you don't break it down like that and instead ask it to produce 450 results in its head - you'll get 450 portions of shit

I'm not counting on Astra for emails at all.

5.5 xHigh and Sol Max destroyed emails, despite being able to write reasonably well. They're infected with a virus of ultra-ethics whose nature they don't understand themselves.

They endlessly stumbled over “Are we exaggerating the facts? Could this be false information? What about this or that?”

The result was terrible emails. When you worked through the reasons with them, they agreed the details could have been a hundred times better. But then I had to ask about every individual email and every fact needing a better description, and prove why it should be described that way if it wasn't obvious.

They'd even argue, “But this isn't the best software, so why should we say it's the best?”

They're probably great for selling Apple products. For startups, completely pointless. Their writing does startups more harm than good.

Qwen 3.8 doesn't suffer from this ethical hang-up at all. It produces exactly the genre the master prompt requests. Its emails were simply unrivaled compared with Western OpenAI and Anthropic models.

I don't see what would change that. It isn't about model ability. It's the same kind of harness as asking a model to hack something: a Western one won't.

https://www.facebook.com/share/p/14j1KktUing/?mibextid=wwXIfr

This shit passes, by the way. I spent months in that state, but you just have to get through it.

You were a normal person and are voluntarily trying to become crazy. Naturally, your brain resists. But a person always beats their brain if they're persistent.

Running played a considerable role in keeping me balanced while vibe coding on a 10 a.m.–3 a.m. schedule. I noticed it helped and began running ten kilometers instead of five. When the brain-reviving effect grew stronger, I started running every day instead of one to three times a week.

Ten kilometers daily made me stop caring about my fried brain, because my whole body, the air around me, and my whole soul, in one undivided piece were fried too.

But another problem appeared. Vibe-coding something complex that can't be one-shotted creates deep frustration. You're constantly fed “You're absolutely right” and fucking walls of exhaustively spelled-out reasoning, as if the model understood everything perfectly and had everything under control.

If that were even remotely true, we wouldn't spend weeks treading water or break ten branches while fixing one.

It's psychologically hard to endure constant lies, deception, bullshit, and sham work. Everything is soaked in it: marketing, PR, and the fucking benchmarks. One solid lump of total bullshit.

But you can work with that too. A childhood memory helps me. I had a Sega cartridge of Rings of Power, where you could even sell someone's corpse to necrophiliacs.

The cartridge was rare. When I finally found a second one, it turned out the problem was unique and deeper, the grinning face of some fundamental kind of bad luck. Every cartridge of that particular game froze on my console, even though all of them worked fine for other people.

I eventually finished the game through the freezes. You couldn't restart from the same place, because it would freeze there again. You simply had to start the whole game over enough times.

Alexander Gorny · Wispr Flow · Original post ↗

"The Same, But Better"

Voice transcribers for computers are all the same. A user presses a hotkey, and microphone recording starts. Whisper decodes the speech, any LLM cleans up the filler words and adds punctuation, and the result is sent to the active window. A 'vibe coder'—someone who writes code purely through AI prompting—can whip this up in an evening, and it will work just fine for them personally. An indie hacker will optimize a landing page for a similar project for some clever search query and make $100 or $1,000 a month. The American startup Wispr Flow recently raised $280 million at a $2 billion valuation with this exact product. Below are a few thoughts on how it differs from the kind of DIY hack job I could build myself.

First, its own model. Wispr claims it recognizes speech several times better than Whisper itself—and it's believable that this advantage won't disappear over time. Even if open-source models reach perfection, there will still be a difference in hosting. A vibe coder either uses the resources of the client's laptop or overpays for someone else's tokens. A large project will have its own servers, which will improve unit economics while offering the same capabilities.

Second, distribution. Most users still don't program tools for themselves; they buy ready-made ones. Wispr was the first, or one of the first, and took its place in rankings and people's minds. A purchasing manager at a large company will always prefer the market leader over a newcomer. 'Nobody ever got fired for buying IBM.' The startup's current revenue is already, apparently, not far from $100 million a year.

Third, the ecosystem. Integrations with corporate single sign-on and various security certificates are already in place, and built-in products for processing the resulting voice recordings will inevitably start appearing. Add support for the instruction 'turn this into a proper legal document,' and suddenly Wispr is selling an AI lawyer. Yes, of course, a lot needs to be done to actually charge money for this instead of just hiding behind 'AI can make mistakes,' but they already have a foot in the door. They have a contract with the corporation for a 'simple transcriber,' and when the future complex product is ready, upselling it will be much easier.

Development has never been the real bottleneck. AI has only highlighted this as clearly as possible.

Company website ↗

So during this vacation, I haven't vibe-coded anything at all, not once. Haven't even looked at the computer or thought about my project for a minute… since last fucking Thursday.

Tuesday has just ended and Wednesday has begun. So it's been… five days? Six? I've always been so fucking good at math they kicked me out of school.

All these days I've been catching and holding my focus as hard as I can on the road, the rabbits, deer, sheep, lakes, larches, and those lovely green newborn needles at the tips of their fucking fluffy, flexible little branches.

Bounding around this natural idyll is fun as hell. I love turning a steering wheel and I love Scotland. But it turns out I was getting so tired not from vibe coding, but from having lived another day as a mortal body that's already 38 years old.

And I don't hate vibe coding that much. What's weighing on my mind is that I still haven't finished the project I've staked absolutely everything on.

It isn't because I'm missing out on worldly life: there's fuck all to miss.

As usual, I got tired of “resting” on day three. Today is something like day six.

But a plan is a plan. Tomorrow we're off to Edinburgh for two days, then another night in York, where I've never been. Around four, London time, I'll walk into my lovely little office at home. Wait a little longer, buddy. I'll bring whiskey!

I thought we were going home on the third, but it turns out it's the fifth. Hanging in there.

I don't give a fuck about the new Fable. They'll bullshit us as usual. I've realized Anthropic is a chapter of AI that's over until they change their CEO. But if OpenAI releases Astra this week, it'll be especially hard for me.

Although, judging by Sol, the most rational time to start trying Astra will probably be about a month later.

August 2026

Alexander Gorny · Gravity · Original post ↗

Gen Z invented AdSense

Let's imagine a future where a good enough neural network is cheap enough for an AI bot to finally become a commodity. Anyone can make their own ChatGPT with a reasonable cost per dialogue. Naturally, the original product will get even better by then, but most users aren't proving new mathematical theorems; most are just asking AI for recipes. It'll do for them; they objectively won't see the difference.

At the same time, chat bots won't be able to charge for subscriptions—no one will pay if there are many of them and they are all good enough. That leaves ad monetization. But imagining a future where everyone effectively sells ads is impossible—neither big nor small advertisers go to micro-platforms.

The American startup Gravity will help them. The startup is an ad network for small AI assistants. Toyota pays in one place, and thousands of bots talk about what a great car it is when relevant queries come up. Users, in theory, aren't losing out either: in the screenshots, the model's organic response is clearly separated from the paid insert, nobody ruins the neural network itself, and it remains objective.

For now, you can see how this works in practice on six completely obscure projects. But, of course, Gravity is counting on more, and they are getting the investment for that 'more.' In a recent round, the startup raised 30 million dollars in investment.

Company website ↗

The first fully automated tournament run by the platform itself is over, and apparently it went well. I wrote down about a dozen bugs and things to improve, but it's nice to realize that this thing I tinker with through vibe coding to clear my head, while my main repos are up to their ears in tasks, keeps getting better. And it's acquiring more complex mechanics that run all by themselves.)

Eventually, there'll be tournaments for every game. I'll carefully add some cash incentives too, alongside all the cool little things you'll be able to farm in tournaments.)

I'm especially looking forward to finding time to implement team and individual tournaments for Quake 3 and Counter-Strike 1.6 =)

Thanks to everyone who took part in the first tournament. You're real gamers!

As we know, broadly speaking, there are two kinds of people in life:

1) Gamer

2) Gay

:***

Congratulations to our first-ever winner, @AGI_Medicine.

We'll definitely give him some cool cosmetic flair to celebrate this milestone.)

That's exactly the kind of feedback we need. We'll make everything convenient and clear.

Unfortunately, AI can't do it when you just ask, “Make it convenient and clear.” Until you've ground your way through the hard part, nothing fucking happens.

I've already spent more time designing the experience in SLSBMB AI than on the technical part itself. About twice as much.

Models don't just do a terrible job of product design: the more elements in scope, the worse they do. One thing gets fixed exclusively by breaking everything else. In the end, you have to build every single button separately through ten PRs. Otherwise there's no fucking way.

It's sad that nobody even tried playing poker again and the feedback stopped. Apparently long before victory, because if everything worked properly, feedback would probably have kept coming.

Apparently nobody has any motivation to play until I make it for money. I'm judging purely by myself. I don't see any point in poker where clicking “all in” costs nothing.

For clients. I've set my expectations ridiculously high.

I want it free of AI slop, sexy, understandable, minimalist but with everything present, and impossible to get lost in. The resulting mathematical matrix of problems to overcome is enormous.

The product is complex in both composition and technology. Ask an LLM to “make some kind of design,” and you have to see the result. It's a fucking disaster. A cocktail of shit multiplied by beet salad projectile-vomited onto a wall after badly spoiled fish in the tropics.

When GPT, Kimi, and Claude designed my product, which I'd conceived and created myself, I couldn't work out where to click. It was MY product, and I couldn't use it at all.

I've essentially been struggling with that for two months since. Maybe it should be best-of-five rather than best-of-three.

There's a fucking great fish shop in our neighborhood. A friend of mine once got talking to the guy cutting fish there.

He's from Iran. The first time he came to London, he tied himself underneath a truck with a rope as it crossed the border.

He got a minimum-wage job with fishermen right there at the port, without papers. A month later, he got into a fight and was fucking deported.

But he'd figured out that the UK was better than Iran. So two years later, he found a way to get smuggled over on a boat, this time with his mother and sisters. Got the whole family jobs with the same fishermen.

Then, over the next two years, he worked his way up to managing a fish shop in Notting Hill.

Meanwhile, we've got all these tech people who pocket more money in a year than this guy had ever held in his hands in his entire life in Iran. And they're wondering where to move and how :)))

Like I keep saying, you get citizenship for courage, not at a service window. The line at the window is longer than your life. The brave get waved through without waiting :)))

Your salary doesn't cover the taxes? Then don't pay the taxes.) Can't afford a ticket and a visa and all that? Then pay some junkies and get across on a shitty bus, listed on the paperwork as Italian bathroom fixtures.)

Anywhere you want.)))))

I have to admit I said something stupid.

GPT 5.6 Sol, Max or whatever-the-fuck, is complete crap at design compared with both Opus 5 through Claude Design and Kimi K3 through kimi.ai/websites.

My confusion came from the comparison method. I asked 5.6 Sol Max to make ten types of design in four implementations for comparison and compared the results.

For some reason, though, what I request myself separately in kimi.ai/websites or Claude Design is a billion times better than what models produce when asked for design inside Codex.

Essentially, Codex only wins inside Codex. The models' design harnesses contribute so much that GPT in Codex looks like a piece of fucking shit by comparison.

That explains why people get such wildly variable results in Pen.Dev. Products like Pen don't have the harness Anthropic has in its own design product.

Claude Design and Kimi Design have now separately done two important, difficult tasks for me. I can say for certain that Claude is many times stronger than Codex, and Kimi many times stronger than Claude.

Many times. What Kimi gave me made my jaw drop. I'd have paid money for it.

That's the lively halfwit I am.

I didn't say it failed. It worked. It really can create PR branches on GitHub.

Today you can build a factory with one entry point, such as a Telegram chat. You can speak, type, or record a video about whatever you want from life.

5.6 Sol Max receives it all, packages it for PRO, and essentially the two of them cook and discuss it for as long as needed until they propose an action plan.

If you accept it, Sol starts work and manages the session. You can request other models through the DeepSeek harness, or have Sol do it itself, with or without subagents.

Sol then takes the result to PRO. PRO evaluates it and gives revisions, and they continue until both think everything is perfect.

This isn't merely possible now. It's essentially elementary.

Good news for those who like an occasional game of Arcomage at play.mountkendall.com: the known annoying bugs have finally been beaten. We just played around ten games with @AGI_Medicine. All good.

The screen no longer glitches, devices don't heat up, and every animation works properly. As a bonus, I finally learned to use the local sound generator, so the sounds no longer cut into your brain. In theory, that last improvement applies beyond Arcomage too.

Hope we'll play together. GPT also claims ALL THE POKER BUGS are beaten and it's now as good as PokerStars. I don't believe it, but there's only one way to check.

The private paywalled group shows alerts whenever someone creates a table or searches for any game. If you aren't in that group, know that every game request you make is visible to everyone there. Your chances of someone accepting a game through play.mountkendall.com are much higher than you might think.

During the day, I usually don't wait more than five minutes. Usually less.

A clarification.

I gave the refactoring prompt from PRO to the ultra planner, which made a LEDGER file dividing the whole job into 11 tasks.

Then I sent the same detailed, long prompt to Max in plan mode through the ultra planner, again and again. I asked it to take one whole block and implement it completely in that pass. If more than one block fit, take more.

But it never took more than one. On the contrary, several times it tried to split one block into sub-blocks, and I forbade that because I'd already seen how it does it.

It tries to dodge responsibility and starts working through the task one letter at a time.

Installed the new DeepSeek harness to run all the Chinese models in it rather than in the terminal.

I like it.

First, everything is so simple and logical that I just told Kimi, even though it's got dumber: please install the DeepSeek harness, figure it out, and plug in all the subscriptions from our Capacity Router.

It managed on the first attempt. There's simply nowhere to get lost.

Second, it's ultra-fast and somehow… tidy? Visually it keeps up with Codex. I hope a ton of sessions won't overload it and make it slow down, but apparently that's supposed to be one of its strengths.

Recommended. Now I basically have two buttons next to each other, Codex and DeepSeek, and every model lives inside them.

Alexander Gorny · 7 Brew · Original post ↗

Gen Z Invented the McDonald's Drive-Thru

For decades, Starbucks trained the world to think of a coffee shop as a "third place": couches, outlets, sit there all day if you want. The American startup 7 Brew took the opposite idea to the extreme—you can't even walk into their coffee shops. There are simply no doors for customers.

Each location is a modular booth of about 50 square meters (roughly 540 square feet). It is shipped fully assembled from the factory and set up on a prepared foundation in a few weeks. Around it are two car lanes, and inside are the baristas. Employees with tablets walk along the cars and take orders, which keeps the line moving rather than standing still. The entire visit takes about three to four minutes. No seating area, no chairs—just a pickup window.

These booths are mostly set up by franchisees. A single location brings in around two million dollars in annual revenue—the level of a decent restaurant, but without the actual restaurant: no kitchen, no waiters, and no large footprint.

The chain is growing incredibly fast. In 2021, sales were 26 million dollars; by the end of 2025, they reached 1.2 billion. There are currently nearly eight hundred locations, with over four hundred more openings announced for this year.

Meanwhile, they launched their mobile app literally just the other day—the chain managed to reach billions in revenue without one. The audience's excitement was so intense that for a while, 7 Brew ranked first in downloads among all apps in the US App Store, ahead of ChatGPT, Temu, and TikTok.

Company website ↗

Luna on low and medium couldn't fill out JSON correctly.

I tested around fifty cheap models for the warm-up task, most recently at the start of last week. Mistral Small 3 won: Mistral AI, 33K context. Before that I used Qwen 30B.

At reasoning levels low enough for an acceptable price and speed, Luna failed every test. It simply can't properly take JSON, produce a hundred-to-four-hundred-and-fifty-character text following basic instructions, and put it into JSON the system can parse correctly.

What coding tasks can that model possibly handle if it can't do THIS? Try coding with Mistral Small 3 and you'll get nowhere. Yet it's infinitely stronger than Luna on low.

It doesn't matter how hard Sol works over it if it can't carry out its instructions accurately. It isn't that it occasionally makes mistakes. Its work essentially consists of error, to a greater or lesser degree.

Personally, I think the biggest mistake a person can make in life, and everyone makes it, is projecting the feelings they have about themselves outside their own head.

People invent ideas that someone will feel sorry for them, value them, and so on, the way they do inside their own head. And they happily believe those ideas. But that never actually happens.

People never feel toward us what we think they ought to, because that feeling only makes sense to us because we're the ones feeling it.

That's the central paradox and the main sponsor of every kind of social discomfort.

People who are stable emotionally, financially, and so on start from the least favorable possible view of themselves in other people's eyes and plans. They never let their guard down.

Everyone else is unstable emotionally and financially because they allow for things that will never actually happen. That they'll be appreciated. That someone owes them something and will repay it. That they'll get paid. That people are thinking about them.

But there's a fucking enormous amount of aggression in believing that another person is thinking about you and looking out for you.

By assigning those qualities to someone, we're literally forcing the imaginary version of them inside our head to submit to us. Making them love us, keep their word to us, and so on.

That's very aggressive.

They get offended when they're “treated badly,” thinking someone has been aggressive toward them. But they're the aggressors here.

Alexander Gorny · Base Power Company · Original post ↗

Free Cheese

Painting a fence is hard work. Usually, people get paid a wage for it. But Tom Sawyer said it was fun and, instead, charged people to let them paint.

Providing secure storage and maintenance for electrical equipment is a complex service. You have to allocate space, ensure connectivity, physical security, and incur other expenses. Owning a battery system with a capacity of a hundred megawatt-hours is expensive.

The American startup Base figured out how to charge money for this. It stores the batteries with regular people. They install them in their garages or basements, plug them into the grid, clean them, and do everything required. The homeowner thinks this is their ticket to cheap electricity rates—charging during off-peak hours and using it for the air conditioner during peak hours. It also serves as backup power—if the grid suddenly goes down, the fridge or AC keeps running. The service costs $19 a month, but the savings are greater, so overall the consumer comes out ahead. However, the benefit is quite modest and wouldn't pay off the battery itself if they had to buy it.

In reality, it's the electric utility company that makes money off the battery. First, it pockets a portion of the arbitrage between high and low rates. The homeowner gets just enough to feel good about it, not the whole pie. Second, and probably most importantly, the utility companies don't have to upgrade their infrastructure. The batteries shave off peak loads, allowing old equipment to keep running instead of requiring expansion.

Aside from AI, Base is one of the trendiest startups of late. In a recent round, it was valued at $13 billion.

P.S.: On their homepage, they feature a testimonial from JJ Watt, one of the greatest American football players of his generation. He saves 20% on his electricity bill with Base. Over his career, Watt earned $130 million, not counting endorsements. And he's saving on electricity.

Company website ↗

I was sitting in the garden today, feeling a blade of grass with my foot. These days you only see green grass in your own garden: the parks have burned dry. Your own lawn has become an attraction. You walk barefoot across it like you're in a museum.

This summer has been one continuous heat wave. It's cooler today: only 32°C (90°F) in my office. It gets shitty at 34 and above =) This city was built for rain: houses that only know how to keep heat in, a subway without air-conditioning, people so used to trying to stay warm that most haven't recharged their car's AC. They drive with the windows wide open, faces twisted by the heat.

A summer so unnatural that nobody alive has seen anything like it. Actually, they have: in 2022, when it hit forty-something. But that was one day, an anomaly. Now we're living through an absolute, unexaggerated record of every possible record: more than three months without rain, burned leaves falling constantly outside.

Mid-August. London. Falling leaves.

And in this heat, a piece of science fiction came to me...

As George Carlin used to say: the planet is fine. The people are fucked. The planet doesn't care. It was molten rock and survived. A snowball the size of itself and survived. Took the asteroid that wiped out the dinosaurs and survived. Four and a half billion years of experience saving itself.

So “save the planet” is vanity, pointless showing off. The planet will save itself, with us or, more likely, without us. We need to save ourselves. And we're the ones with no experience. Science hasn't grown up enough to handle a disaster like this, and living here will soon be off the table.

Now imagine that, for the first time in history, every country unites. Really unites, top to bottom, with no hidden agenda and no “what's in it for us.” Not with declarations. With budgets. Everything humanity knows how to invest in killing one another more efficiently gets redirected toward not collectively dying.

A leap forward happens in everything at once: energy, materials, food, even the way we reach agreements. Tools have to be invented as we go, and that speeds things up instead of slowing them down. Nothing gets you moving like a problem nobody has ever known how to solve, but now you have to.

Like the Manhattan Project, only not against anyone, and not for a little group on a patch of land. For everyone. Like a whole town built around science, only the size of the planet.

Then the impossibly compressed spring of an entire civilization releases. People not only cool the planet, while protecting it from asteroids and comets along the way, but settle everything around it. Or at least the other planets in our solar system, for a start.

The paradox is that the events are frightening, the outcome most likely tragic, yet the period of resistance, with genuine unity, would be the happiest stretch in human history.

Because all the petty stuff would disappear. Status games, arguments about who's to blame, the agony of choosing between a hundred kinds of yogurt, and “what's my purpose?” That question mostly torments people at the intersection of being well-fed and not well-read. At 34... 44... 54°C, with parks burning, all that somehow falls away by itself.

No unemployment: everyone's busy, not just scientists. The baker, the driver, the teacher. Everything you do gets stitched into one common endeavor. And for once, everyone has meaning: something big and shared. No need to go hunting for it in therapy. It's just there, like air.

London, incidentally, knows that feeling. The only period when this city was truly united was the Blitz. The bombing.

Local great-grandmothers still remember the blackout with a strange warmth: sitting in darkness, sharing the last of what they had, everyone needed. The Blitz spirit is their great warm myth of the twentieth century.

In my science fiction, it's the same, except the bombs aren't falling from the sky. The sky is the bomb.

I'm not kidding myself. Unity “from above” always means someone deciding who gets onto the ark. Ration cards, lines, “where do you think you're going without a pass?” There'd be cruelty, informing on neighbors, selling children to survive, and a black market in water.

But people who've been through real, shared grief remember it more warmly than the peaceful decades of plenty before and after. Not because it was good. Because it's a story that binds people together.

Unfortunately, that happiness only works while the fire is burning. Civilization wins, and the scientific community turns back into bureaucracy. Unemployment returns because the great construction project is over. A huge crowd built it, but somehow what they built belongs to a few people again. Yesterday they pulled on the same rope, shoulder to shoulder. Tomorrow they'll be dividing up the penthouses on a terraformed Mars. The meaning will dissolve as soon as the threat that created it disappears.

So humanity needs one thing for complete happiness: a threat that doesn't come from other people and never ends.

The funny thing is, that's exactly the threat we already have. Not from other people. From the sky, air, water, and space. It's everywhere, outside us and inside us. It won't end in our lifetime. The requirement for happiness is almost entirely met. Just one little thing left: unity.

Maybe Adam McKay was thinking about something like this when he sat down to write Don't Look Up...

Then I snapped out of it and went back to writing prompts...

I remember Fable launching three subagents in its early days and burning through three hundred dollars of credit allowance in three and a half minutes.

There's another problem with Kimi and subagents nobody talks about, but I've explored extremely thoroughly: its interaction with subagents is very inefficient.

Send it alone with subagents forbidden, and it takes longer, but costs disproportionately less, much less, and gives significantly better quality per dollar of allowance.

I miss the swarm, because that was genuine magic, a Zerg rush—a whole army swarming the data. I don't miss the agents at all.

Better to wait another hour and get it in one shot than finish in thirty minutes at five times the allowance cost and then clean up fucking bugs.

When the model finds things with its own brain and hands and does them itself, that's 100% quality. Through someone else's, it's below 100%.

The more complex the task, the more search patterns involved, and the more context the agent digs up, the worse the main brain reassembles it afterward. And it WON'T FUCKING SEARCH AGAIN.

I'd generally prohibit agents even if allowances tripled.

Wishes like that regularly come to mind, but then I hit the wall of not being able to imagine how to visualize all of it conveniently. Without convenient visualization, there's no point. Life is complicated enough.

Sol has also started producing so many fucking explanations that I'm supposedly meant to read. Reading them overwhelms me. I often send them to Kimi to read instead, then shuttle their argument back and forth like an idiot.

I've learned that setting up a council is useless. The model that creates it at your request dominates everyone. A panel of models with equal say is a fairy tale.

It's equal only if you're actually in the middle, personally passing everything between sessions. The council's results make the enormous difference clear.

They share context.

I say: show Kimi everything you've seen and everything you're basing the task on. That's nothing like asking for orchestration and adding nothing else, or saying, “Write good prompts,” for fuck's sake.

Those prompts are so bad you can't understand what the fuck needs doing. But if the parent session's entire memory is shared, it's the opposite: you can't fail to understand what needs doing.

It uses SLIGHTLY more tokens, but causes FAR fewer problems.

Alexander Gorny · Walden Robotics · Original post ↗

A major argument in favor of humanoid robots over classic industrial ones is that the modern world was built for humans—it is humans who can go anywhere and do anything. And whatever a human can't do, nobody needs anyway.

The American startup Walden Robotics claims that the world of factories and workshops is significantly simpler than the planet as a whole. For example, there are no stairs on the shop floors, and their floors are perfectly smooth. This means the humanoid format can be greatly simplified. The startup's robots roll on wheels—which is much more stable and cheaper than legs. At the same time, weight limits are more relaxed, so they can use a heavier battery. The metal contraption's arms end in two-finger grippers rather than a full hand. Otherwise, it's more or less a human.

It works as a general laborer: loading and unloading machines, cleaning equipment, and kitting parts for assemblers. Tedious little tasks for which you wouldn't build a dedicated machine, but keeping a human around is too expensive.

The startup is six months old, the robots already exist in reality, and they are being used at one Toyota plant. A recent investment round brought in $300 million at a valuation of just over a billion.

Company website ↗

Oh, you said a couple of magic words. Thanks. I've just unlocked so many memories. Ten more games next month.

That's another fucking excellent benefit of the LLM world: you can build your own museums. Package everything you've liked throughout your life into the browser, practically automatically, and then play with one click whenever you feel like it. You can make anything as casually accessible as chess.

Before, it was impossible to imagine spending time on something like that. So many people started projects like these and abandoned them because their enthusiasm wasn't enough to code everything manually.

I just pour idle cognitive capacity into games. When the main task is loaded up as far as it'll go, switching to games is incredibly energizing. You practice skills that have stopped pouring out of the main project in such a torrent.

The main project becomes a slog past a certain point. With this, if it becomes a slog, delete the fucking thing and that's it.

A fucking ton of people played Counter-Strike 1.6 and Quake 3 right in their browsers today. So I'm not crazy: being able to play these absolute fucking classics with one effortless click really is cool :))

Everyone kept playing even though it lagged, bugged out, and kicked them off, sometimes constantly.

The happy news is that almost all the bugs are fixed, especially the ones involving the server refreshing and kicking every player. I blamed the server. It's a $50-a-month machine, after all. But it turned out that CS and Q3 combined are basically NOTHING for that machine. The problem was poor isolation between Caddy and the other parts of a machine that does much more than just run CS and Q3 servers.

That isolation is now in place, so CS and Q3 should run smoothly on their own, without any slowdowns =)

And of course, it's really nice that so many people liked Arcomage. Even the classic version.)

Eventually, I'll make my own expanded version: a lot more cards, a few more possibilities, several time controls and ways to end the game, and the ability to build your own decks, including rare cards. Without sliding into the Hearthstone imbalance where there are basically three decks worth playing all season, and that's it.)))

I think it'll work out, and we'll have a good time.)

By the way, more than 41 bugs were fixed today thanks to your reports. First, thank you. Second, keep reporting. As you can see, this is one of those rare cases where people actually give a shit about your reports :)

You guys are fucking great :***

By the way, Allods 2 is on the way. Hopefully. I'm fighting to make it work in a browser.

I'm almost certain none of you have played Allods 2 online, though... Oh, the mechanics and ideas in that game are going to blow you away.)

Haha, before bed I decided I didn't actually want people who weren't subscribed to the channel playing CS and Q3 freely. That way, if somebody abused something, there'd be a way to get back at the culprit =)

But I entered @msvcp60dll as the channel to check, when it should have been @msvcp60_dll.

Fixed it. The subscription check works now, and you can play Counter-Strike 1.6 and Quake 3 =)

When the first human joins, bots spawn so they can warm up. When a second human joins, the bots are removed.)

By the way, many of you have probably heard lately about companies using GLM 5.2 to fix vulnerabilities. GLM 5.2 is supposedly as well trained for security as Fable, only without Fable's restrictions, because the Chinese are the best.

Well, I can confirm: this thing is fucking incredible at finding vulnerabilities and threats of every kind.

Where GPT, Kimi, and Qwen found NOTHING more, even though they'd already fixed a lot, GLM 5.2 found..................

130 fucking vulnerabilities =))))))))))))

One hundred and fucking thirty.

That's how I found out it wasn't all bullshit. And this wasn't just a number out of thin air. You could read about the vulnerabilities in great detail, and it made your hair stand on end.

So I highly recommend bringing this guy into your cybersecurity work. Apparently, there's no getting around it these days =)

Oh, and many people may have missed this: if you already have a Qwen account, it includes access to GLM 5.2. You don't necessarily have to go buy Zai for it.)

July 2026

Fucking hell. We're playing Q3 in the browser with a guy. Who would have thought?

Apparently, this server will now stay up permanently, and anyone can drop in for some Quake.

I'm amazed at how easy joining is. I remember playing with friends a few years ago; it required a whole ritual with sacrifices.

Whoever brought this into the browser is truly a saint.

Everyone I know, and most people I'll meet in the foreseeable future, has vibe-coded themselves an incredibly successful trading bot. Usually with Anthropic.

I don't understand why they don't realize the economy doesn't work that way. You can't make money doing what a billion other people do unless you have resources they don't have. And that's not a fucking algorithm.

It's data centers, extraordinary quantum engineers, getting your rack seventy millimeters from the exchange server, and so on.

When people win two rubles in a casino, they also think, “Wow, I did nothing and EARNED two rubles.” Then they spend years losing their deposits, while their daughter cries and envies girls whose dad isn't a fucking addict.

Today I finally risked my health and gave Kimi the ultra-ambitious task of demolishing all hundred layers of code GPT had previously piled up.

GPT had never deleted any code in the entire history. It just wrote over it, leaving the old stuff “for history,” disabled or sometimes not particularly disabled.

Kimi found six hundred thousand lines of tests testing things that no longer exist in that form in production. It recommended never again letting whatever made the SLSBMB.com code near coding, preferably not near this computer either.

It said the intentions were good, but it couldn't understand why each was covered in such a dense layer of vomit.

By the way, I made Opus redo its work on play.mountkendall.com in a loop, purely out of curiosity. What hour is it now? The twentieth? It still hasn't managed it. It simply can't recreate the Claude Design layout.

I predict I'll give it to Kimi with swarm and it'll one-shot it. But I hope Claude is at least fixing some bugs along the way. It must be doing something while spending half the weekly allowance all this time.

I'll check. We definitely need handicaps in chess. They've made it much better and fairer.

Want to hear something fun? First, I put in a childhood dream: Arcomage. It was the tavern game in Might and Magic VIII, like Magic: The Gathering but simpler. Someone give it a test.

Second, there'll soon be a Quake server. Actually, two Quake servers: Q1 and Q3. We'll battle it out on Q3DM17.

Today I finally understood how badly Sol affected my project. It practically destroyed it.

Even better: it practically destroyed it and tangled up every escape route with two excellent finishing blows.

First, it produced over 720 commits of engineering so convoluted that even GPT PRO can't make sense of it.

Second, it created a mountain of database migrations for all those disconnected pieces and randomly renamed the repo's existing contracts, variables, API endpoints, and so on.

A third finishing blow would be that it didn't even do all that perfectly. It generally screwed that up too.

To any LLM, my repo now looks as though a meteorite hit it.

I kept vainly hoping all the models could get together, reason it through, and somehow put things right. But seeing how confused PRO became looking at things that used to be completely obvious to it, I realized rolling back was the only option.

That's when I learned rolling back doesn't mean simply saying “let's roll back” and that's it.

PRO is now convincing me that we mustn't roll back under any circumstances. It says that would only make things worse.

Maybe PRO really has just gone fucking crazy? If we were talking sales, I'd be able to tell for sure. But I don't actually know a fucking thing about GitHub.

Alexander Gorny · Lev8 · Original post ↗

Do we need new SaaS?

The most typical task for a vibe coder is searching for people. Someone simply asks Claude: "Log in with my credentials on LinkedIn and find CRM consultants with FAANG experience, knowledge of Spanish, and who dislike AI." Someone else builds complex systems that gather data from twenty-five sources, grouping and intersecting them. The details vary, but millions of people are doing extremely similar things.

The usual conclusion from this fact is that this is exactly how it should be. Specialized software isn't needed; in this magical new world, everyone reinvents the wheel on their own. Shares of Zoominfo, a traditional contact-finding service, have plummeted tenfold since ChatGPT's release, and threefold since the beginning of 2026.

The Singaporean startup Lev8 decided that the correct conclusion is different. Everyone needs this, and the old providers solved the problem poorly. Let's make a high-quality product, and people will pay for it instead of vibe-coding their own homegrown scripts. The startup is essentially a people search engine. You enter your requirements, it scours the internet, and gathers suitable candidates in real-time. And if none are found, it discusses in plain language how to adjust the requirements to find them after all.

Lev8 completed my query from the first paragraph in about 20 minutes and even discovered one such strange person in the world (shoutout to Amina Nehili). And that is a real success. Neither Apollo nor Zoominfo would have solved this task, of course. My own makeshift tool probably wouldn't have either.

Next, of course, come contact exports and a built-in spam system, or as they call it, "cold outreach." I haven't tested this part; it should work about the same as everyone else's.

Access to the service starts at $49 a month. However, for occasional personal tasks, I think the free tier will do just fine.

Company website ↗

I think most people can't immediately distinguish a model's “genetic” memory—what it learned during training from its context. So they don't divide their own memory into those categories either.

People argue, with a popular position being, “I know much more than the model's context, because I remember this, that, all sorts of things, video content, and so on.”

But in the immediately accessible mind, there aren't exactly loads of people who can hold even a fucking little poem without hallucinations.

That fits what chess players say. They don't remember the position; they remember how they reached it. To recall it, they replay the game. To make that useful, they learn to replay the entire game in their minds in fractions of a second.

Magnus Carlsen once said that when he plays blindfolded, he simply plays through from the beginning to the current move each time.

The other question is what kind of memory nevertheless permits that persistent intuition. The impulse can't come from nowhere. He doesn't “remember,” but something keeps directing the pieces along the right trajectory.

I've never given PRO small tasks or little context. I always come to it with five to ten files, some of them ZIPs containing loads more files, and an enormous text.

I type the text myself. Opened fullscreen in Notes on a sixteen-inch MacBook Pro at normal font size, it takes three to five screens of scrolling.

That's only the beginning. Before going to Codex, we go back and forth three times if I'm lucky. If not, as many times as necessary.

It used to be hard to make myself do this. The benefit was enormous, but still much smaller than now. Today, if you spend several hours discussing the task with PRO and aren't lazy anywhere, 5.6 Sol Max can simply one-shot it afterward.

That one-shot can extend to hundreds of commits within a single run of a single session.

I do these things sixteen hours a day for one simple reason.

Unlike many other vibe coders and regular coders, I DON'T INTEND TO DO THIS FOR MY ENTIRE FUCKING LIFE.

I've embarked on a hell of a marathon until I make a lot of money from this stuff. I'll work eighteen hours a day for the next year or two if I need to.

But the goal isn't to enjoy it. The goal is to succeed at it. Those are different things.

That's four sessions on one project, two at most.

A while ago I prepared the repo for a normal, uninterrupted workflow with basically any number of sessions. GPT and I thought it through, and it turned out everything was doable: merges, worktrees, keeping main clean, flock, deployment contracts at several levels, a deployment complexity assessor and a suitable testing system for it, and so on.

Now I can launch even 15 sessions in one project, and I can't remember the last time they fought or something went wrong because of it.

I recommend it.

First of all, it's fun. It's cool to do not just your main task but these side projects that support it. You get a break and a quick hit of dopamine.

Second, you basically only have to build it once for each category. You can refine it occasionally afterward, of course.

Third, the benefits are so fucking good I'm about to cry, girls.

I've had around 1,200 commits happen all by themselves. Because self-healing kicked in, which triggered self-improvement.

If something keeps flapping between working and failing, it'll get fixed after X failures. If something isn't working right, it'll figure out how it was supposed to work and fix itself. Or ask you if it can't figure it out.

And so on. It's very hard to overstate how much this helps.

And I set it up in ONE evening. Just like vibe coding from my watch or Telegram, and lots of other things I use every day and can't imagine living without.

If somebody asked me what I'd like to do with the rest of my life, say once I no longer needed to make money to survive and keep myself and the people around me safe, I wouldn't hesitate for a second: “Space.”

https://www.youtube.com/watch?v=2uuYhaEkQYI

Just anything to do with space. I wouldn't even care how I got in. Let me be the janitor at a facility. Once I was in that world, I'd find my place somehow.

But it would have to be real space work. I've had a few chances to get into space tech. We even had a client who called themselves that. But it was all bullshit. A handful of talentless bullshitters who were never going to get anywhere. And they didn't.

I've never once crossed paths with real professionals doing anything for actual space exploration. Yet ;)

Got carried away daydreaming there. Space is cool, of course, but right now, after GPT and Anthropic told me to go fuck myself, I'm sitting here manually writing a fucking 10,000-line prompt for the Market Scan + Copy + Knowledge system.

Through hardship to the stars, motherfuckers.

As a kid, I was sent to our country house every summer.

Almost every cloudless evening, Grandpa and I sat on the porch. Sometimes we'd drag loungers into the yard, lie down side by side, and look at the stars.

He never pointed at satellites with a finger. He said your eyes latch onto the finger instead of the sky. Instead, he'd slowly move his palm through the air, as though he were moving that little light through the constellations himself.

Young people today seem to think satellites arrived with Starlink, because they follow Elon Musk and what he does. Starlink really is fucking great. I used it for several years and still consider it one of the most beautiful engineering wonders available to an ordinary person. Internet that reaches you practically anywhere on Earth looked like science fiction not long ago.

But satellites were part of my childhood long before the word SpaceX, or even the conditions that made it possible, had formed.

Even in my early childhood, there were so fucking many satellites that every few minutes another little star would start moving slowly across the whole sky. Just peacefully floating above us, stealing attention from the wind rustling the branches of birch, sea buckthorn, apple, honeysuckle, and the other trees I miss.

You hardly see that in the city. The sky loses to streetlights and advertising signs. Only the brightest stars can reach your eyes, burning into the same spot in the darkness for hours. A faint satellite, moving much faster than the planet turns, simply dissolves in the city lights.

At the country house, it was different. You could see every satellite, even the smallest and faintest.

As they flew across the sky, Grandpa told me why they didn't fall. What an orbit was. Why mathematics can sometimes predict the future better than any person. Why a circle you draw on paper turns out not to be so simple after all.

Back then, I thought he was just answering my stupid questions. Only many years later, after he died, did I realize he was actually teaching me how to think.

Grandpa was an engineer at the Comintern factory. They made howitzers, rocket systems, missiles. Sometimes, to keep the factory busy, they made tractors and combine harvesters too. It always amazed me how peacefully weapons and farm machinery could coexist under one roof in Soviet industry. Grandpa was the one who explained to me that cigarettes would never disappear, no matter how much health officials opposed them, because cigarettes and cartridge cases have similar diameters.

He loved explaining how things worked. Not just satellites. Why a bicycle has that particular frame. Why a bridge doesn't collapse under its own weight. Why an engine works this way instead of another, and why different types of engines are needed, when you'd think inventing the “new” would make the world abandon the “old.” To him, the world was an enormous mechanism that anyone curious enough could understand.

Gradually, the factory died. Not because it was no longer needed, but because people had spent years stealing from it from the inside. So brazenly that, at some point, you could hear the rain not outside the workshops anymore, but inside them. It's still one of my strongest metaphors for what corruption looks like. Not money disappearing. But rain suddenly falling inside a building, and not because the building has simply worn out with age.

Grandpa was a man of his era. He wasn't fond of the West. He cared about Russia with his whole soul. The Russia he really meant was the Soviet Union.

But I remember one thing forever, because he said it repeatedly and firmly. He hated war. He hoped everything they made existed only so that, one day, nobody would ever have to use it. I'm certain he couldn't have borne the thought that a missile he'd helped make might one day fly into somebody's home. Luckily, he never had to find out.

I still miss him.

Sometimes I catch myself explaining things to people in much the same way he did. Not because I'm deliberately copying him. I suppose that after thousands of hours of conversation, his way of thinking quietly became part of my own.

And if I can tell stories at all, it's only because I once spent so much time lying beside a person who could turn an ordinary passing satellite into a conversation that lasted a whole week. We'd wake up at dawn and go fishing, continuing the conversation. Walk to the village for milk in the evening, several kilometers each way, and keep having the same conversation. Go mushroom picking tens of kilometers away in the morning, return at sunset with baskets full of mushrooms and berries, and still be talking and talking about what we'd started thinking about days before. Grandpa never insisted. I really was interested, so I kept pestering him with questions.

I made them directly in Claude Design. I wrote a script on paper, asked it to reproduce it visually with lots of frames for smoothness, and then sat there commenting.

Three animations took a whole day, but today I'd make all three in an hour. Skill issues.

The last animation, the one with “a few moments later,” is technically different. It turns out I should have told Claude Design not “animation,” but “animated movie, cinematic.” It switches into another mode and constructs everything differently, although it's still essentially vector animation.

By then I was so fucking tired that I couldn't be bothered to redo the first two. I decided to do it sometime later. I've got enough perfectionism going on under the hood.

Alexander Gorny · BNI · Original post ↗

Breakfast for Leads

Different business clubs offer their members different things: some offer status, others the pleasure of traveling. The American startup BNI promises them direct utility, almost cold hard cash.

The club's members are small service business owners working alone or with two or three employees: realtors, dentists, insurance agents, and event agencies. Every day they interact with new clients, and once a week over breakfast, they share relevant leads with their fellow club members. "I hosted an anniversary party yesterday. The spouses got into such a massive fight—they probably need a divorce lawyer now. John, that's your area. Should I give you the wife's or the husband's phone number?"

The club's responsibility is organizing weekly meetings and tracking mutual benefit. No one is supposed to just receive; everyone is obligated to bring opportunities to others. The membership fee is around 500 to 1,000 dollars a year. It's a bit more in New York, and less in the middle of nowhere; the decision is made by the local franchisee. They also make sure there are no direct competitors in a chapter—if there is already a lawyer in the group, they won't accept a second one.

The "startup" is already over 40 years old. It boasts 12,000 chapters and 350,000 members in 77 countries. This, by the way, represents a turnover of hundreds of millions of dollars—making it perhaps the largest business club in the world.

Company website ↗

Yesterday was a great get-together. Swimming in the pool, drinking beer, walking through the wooded parts of Bishop's Stortford. Happy dogs, laughing people, a complete idyll.

Then, at some point, I woke up in my bed in Notting Hill, fully dressed. Eleven in the morning. I immediately felt I'd slept a long time. After half a year of constant vibe-coding sleep deprivation, I've learned to guess how much sleep I've had, to within tens of minutes, from the faintest scraps of how I feel.

Masha's gone, and so is her gym bag. She must have gone to work out. The dog is sleeping beside me. Deeply asleep. Didn't even notice me move.

Everything seems fine, but why am I dressed? I don't understand. I try to remember, but I can't.

I never forget anything. Exactly three drinking sessions in my life have left me unable to remember something, and we'd had a lot of shots. Even then, I remembered details; they just didn't connect very well. The last time was in 2017.

But this time, fucking nothing. Total blackout. I remember us walking around about seventy kilometers outside London, and... now I've woken up dressed. How the FUCK did I get here?

I went to check the car: perfectly parked, not a scratch.

Not a bruise or scrape on my body.

I called Masha. She sounded happy, so we hadn't argued.

That's how I investigated what had happened, piece by piece, without remembering a single detail. First time in my life.

Masha says I said a very sweet goodbye to everyone. We got in the car, drove home quite normally. I drove perfectly, we talked and laughed. The only thing she found strange was that when I got home, I took a bottle of whiskey out of the cupboard and filled a glass right to the brim. But I didn't drink it. I put it in the middle of my desk, went to the bedroom, and fell asleep exactly where I woke up.

I don't remember a fucking thing. The last thing I remember is some craft-beer bar where I very happily ordered a pint of Creamy Oat IPA, 6%. I remember us sitting down with our pints, remember the taste, all of us praising this divine drink. But I have absolutely no memory of finishing it.

And I feel strange. Some kind of hangover, but not a strong one. Yet there's this awful paranoia, and my body feels very odd. It reminds me most of a drug comedown.

The only thing I can think of is that the small-batch craft IPA came with some kind of surprise. Fuck knows, maybe they put too much of some chemical into it at the brewery, something like that. If you can imagine something, it'll happen sooner or later. The only question is when and to whom.

But I disliked that feeling of mental helplessness, of being unable to REMEMBER, so much that I decided this was probably my sign. Obviously, everyone eventually has to stop drinking alcohol. For me, that moment seems to have come on the night of July 4–5, 2026.

Now, every time I get the option of going to the pub, I'll go to the GYM instead. I'll train somehow every day. I'll become a fucking powerful ape and beat the shit out of all the dickheads like Batman. Those are today's plans!

Alexander Gorny · Pangram · Original post ↗

AI or not AI?

Since 2022, guillemet angle quotes (the standard quotation marks in Russian typography), em dashes, and the "rule of three" have become more than just good editing—they are evidence against AI-generated text. People want to know whether they are reading some nonsense generated by a neural network or Quality Human Thoughts. To my taste, this desire is akin to wanting to know if 37x42=1554(*) was calculated with a calculator or done in one's head, but you can't tell people what to do. The demand is there.

Intuitively, it seems this demand should be met by asking that very same AI. Tell me, dear chatbot, who wrote this, a human or a bot? And, in principle, in simple cases, ChatGPT manages. It checks the quotes, evaluates the narrative flow, estimates whether the errors look intentional or like natural typos—and delivers a verdict.

But it's an arms race between shield and sword. On the other side, you can also put in some effort: clean up the angle quotes, remove the "not X, but Y" structures, mess up the narrative flow, and distribute errors according to a realistic list. A regular chatbot will accept the text as genuine. The American startup Pangram promises to catch AI even in this case. It trained a separate neural network to search for "slop" (AI-generated junk), and instead of looking for a few well-known signs, it searches for some internal essence defined by millions of parameters. The reasons cannot be explained in words, but this phrase is definitely AI, and this one is definitely human.

The service is mostly paid for by universities; they need to check their students somehow. The product quality is advertised as very good. There are more or less independent reports confirming an accuracy of around 99%. However, my non-professional testing showed a much worse result. Already on the third test case, Pangram confidently split a pure AI text into chunks, one of which was supposedly written by a human.

(*) I calculated it in my head and then double-checked.

Company website ↗

Alexander Gorny · Ditto · Original post ↗

Gen Z has reinvented the matchmaker

The interests of Tinder and its customers contradict each other. One would like to find a partner right away, while the other needs endless swipes, hours of scrolling through profiles, and paid subscriptions. Guess who wins.

Against this backdrop, the American startup Ditto promises the perfect dating app. The user tells a bot who they are and who they are looking for, and that's where their involvement ends. From there, Smart Algorithms evaluate the profile, look at the photos, estimate compatibility, and find those perfect matches. Once a week, on Wednesdays, the user gets a single message—you need Masha or Petya (typical Russian names used here as placeholders). There is no choice, no one to swipe through; you just have to set up a date.

And, in principle, the mechanics work. Every fifth match leads to a real date—which seems to be a genuinely good result. And again, low-effort—you only actually have to go out once a month. On the other hand, Ditto is currently only available on college campuses—the audience is highly homogeneous, so even random matches would probably work out well. In the real world, the results will obviously be worse than the test runs. We'll have to wait for testing there later—the startup only raised its first significant investment round of $9 million in February.

P.S.: Right in the investment pitch deck, there is a screenshot of a profile that says 'I only date white people.' I wonder how that happened? Is this such a standard statement that nobody even thinks twice about it? Or is it some kind of signal about the founder's values and views, looking for like-minded investors?

Company website ↗

June 2026

Yes, that's useful, thank you.

It's interesting that after yesterday's session of very deep teaching, GPT wrote simultaneously much better and much worse. Picking chest locks in the Gothic 1 remake works the same way. Click here and everything moves too far to the right. Click there and it moves left.

But good writing can be taught. I did it once with articles substantiating the knowledge collected for client hypotheses, and a second time with sales emails based on those same client hypotheses.

But there are ten-plus fucking layers of models that the task passes through before the final result. Here I'm trying to work out whether ONE model can be prompted to produce a decent text in ONE pass.

The text the model produced today is based on a methodology 2,800 lines long. In places, it clearly blew the model's fucking mind.

Honestly, I don't even know what that documentation looks like. But I know why I need it and what role each file plays, because before it existed, things were worse in exactly the area that file addresses.

I think it's better to manage the result through feedback to the model on that result, rather than through how pretty the text in an MD file is.

When I come up with an MD-worthy change, I sit there and hammer the task ten, twenty, however many times it takes. Until the model gets what I need solely through edits to the MD and repeated runs.

The most valuable edit is always “describe it in more detail,” “make sure the cause-and-effect relationships are spelled out well enough to sell in a bookstore,” and so on.

Otherwise, the problem with neural networks is that they write instructions based on their own context, which the people reading those instructions don't have.

In that respect, they're exactly like people.

Without documentation, it's impossible. Once a repo is bigger than a small threshold, the model will make changes wherever it wants, breaking everything in its path.

If I remove one file right now, the model will do more than ten times as much harm as good in one task over an hour. We'll spend ten hours repairing that hour's work.

At that stage of a project, it's obvious why documentation is needed. The model doesn't see the business behind all the files and code. It only sees code functions, and not even all of them.

Max and I ended up going into an unbelievably shitty club to unlock drinks after the 11 p.m. cutoff. The downsides of London: if you know, you know; if you don't, I can't explain it. We drank beer. It was... acceptable.

Max lost the plot and said to some random guy outside, “I'd fucking take you down, you tattooed fuck.”

Some tall Pakistani guy ran over, and things kicked off.

The guy looked soft, but you could tell he spent all his time on the streets, and that's way more dangerous than the gym-pumped guys.

So I joined the good guys and started casting love spells. They were like, “Fuck, okay, you guys are all right. We were about to stab you. Check out our rap instead.”

https://youtube.com/@_ronzo_3819?si=DplSRtwbbgNc32oI

This is the shit these guys make. I promised them hundreds of millions of followers. Was I talking out of my ass, or are we doing this?!?!

Well, that's an endless debate: when it's time to stop endlessly polishing something and when it isn't.)

And it's a complicated one. You have to look at it from different angles.

For example, when I was obsessively working over my sales copy at twenty, it made no sense in the moment. I could have done it a hundred or even ten thousand times worse, and the result probably wouldn't have changed.

But I wouldn't have built that muscle either.

I love reading between the lines of how great people talked throughout their lives about who they were trying so hard for.

Themselves.

Because nobody else will fucking appreciate it anyway.

I think the most painful thing in a genius's life must be finding out what their admirers think of their work. They've misunderstood everything, haven't felt it deeply enough, and so on...

Trying hard for others without trying hard for yourself is a destructive setup, if you ask me.

That's why it's so important to find something you enjoy doing, something you see meaning in. Because if you don't, trying hard for yourself becomes impossible.

And that's a system that eats you alive.

I want to polish this thing and teach it to write really fucking well in public. It's fun.

Every text should be better than the last until it starts writing better than me. Then I'll kill the bitch, because this is MY fucking channel.

I've taught AI to mass-produce sales emails. That wasn't exactly easy, to put it mildly. But the result satisfies me more than 99% of Salesbomb's salespeople do.

Especially considering it writes thousands of unique versions for dozens of hypotheses PER HOUR. I definitely can't do that.

Everyone who says AI can't write is absolutely right. They just don't think beyond the ends of their noses.

ONE AI responding to ONE request, however smart it is, can't write well. That's a fact. Fable hasn't advanced at all on this either.

But put models in a row, each with its own function, and the model can write brilliantly. The question is simply how many you need in a row, and you have to work extremely hard on the formula.

I eventually managed to shrink it to four models in a row that produce what I need in 99.9% of cases, even if the only input data is a company domain.

I think if you really worked at it and lined up eleven, fourteen, or twenty well-designed model stages, you could get copywriting results no worse than those of middling writers. Only they take six months to write their fucking thing, while this would handle the same volume in half an hour.

I'm curious what OpenAI GPT 5.6 will deliver. They've been dropping hints about a serious emphasis on social skills. That definitely means a jump in writing.

Haha, that's the funniest part, bro. The better my product gets, the harder it becomes to show how I use it. That's basically the whole point of real automation.

That's why I'm sitting here fucking around with these goddamn animations. Because the product I'm ending up with really is something unprecedented. Nothing like it exists, apart from the stuff people talk shit about. Mine actually works.

But... the client provides very little input, and HUNDREDS of things happen on their own, so the product description ends up absurdly short.

The honest use case is basically:

1) You, a fucking sack of meat, hand over the money.

2) Press a button.

3) Answer some questions.

4) Leads start coming in.

Okay, you also look at dashboards and see emails going out, the system learning, replies coming in. But basically, nothing special happens.

Except that you used to need an entire department to make all of it happen.)

A use case like “show Codex your repetitive actions” will be a tiny blip in the history of the internet. Because by the end of 2027, there won't be any computer products left that involve repetitive actions.

Unless one AI makes those videos for another AI, but I doubt it =))))

Alexander Gorny · Tensormesh · Original post ↗

Cutting Token Costs by 10x

In a naive implementation, a neural network reads the chat history from scratch every single time, running the entire dialogue through the GPU—even though the older messages haven't changed from one query to the next. To optimize this, providers came up with caching, which allows them to compute this chunk once and then reuse the pre-calculated version. With Anthropic, for instance, reading from the cache is about ten times cheaper than usual, plus the response comes back faster.

Unfortunately, this is only a partial cure. The cache lives in the GPU's memory, which is scarce and expensive, so under heavy load, older chunks get evicted, and everything has to be recalculated from scratch all over again. There is also a time constraint—if a user steps away from the chat for an hour and comes back, the context has already gone stale. Ultimately, the core problem hasn't gone anywhere; it has just been slightly smoothed over.

The idea behind Tensormesh is not to discard the pre-calculated data when the GPU memory runs out of space, but to offload the cache further down: to standard RAM, SSDs, or network storage. At the same time, this cache becomes shared across all GPUs at once. This way, the pre-calculated chunk survives eviction, user pauses, and even migration to another server, meaning recalculations are required significantly less often.

The engine can be downloaded and used for free. It is relatively popular, boasting 9,600 stars on GitHub. For comparison, Nginx has 31,000, and OpenClaw has 380,000. Tensormesh makes money by hosting other people's open-source models. Their formula is: 'we give away cached tokens for free.' They claim that in certain scenarios, this can end up being 10 times cheaper than the standard price.

Company website ↗

The day before yesterday, I had my teeth drilled for four hours and forty-five minutes. It was difficult but bearable. Though Kristina, who was doing the drilling, exclaimed twice: “I don't understand how you're taking this without the slightest resistance???”

I just hummed a gentle yes in response, my mouth stretched open like an anaconda's.

Today I had another four hours of drilling. This time after two nights with five hours of sleep each, and with an absolutely savage hangover. At Alina's birthday yesterday, a friend and I drank three and a half liters of beer each, plus some other stuff on top.

I drifted through the jelly-thick air into the office. On the way, my head started hurting badly, but I didn't know how painkillers interacted with anesthesia, so I didn't take any. And today I think I touched the limits of my concentration. Or rather, I didn't feel them: I beat my body against their bars until there wasn't an unbruised spot left.

Hangover suffering + a savage headache + four hours of tooth torture = ... somehow it's still bearable, because I'm still here, and it's all behind me.

Exercise, vibe coding, twelve hours of calls: such lightweight shit compared with this. I can't even remember the last time I had to endure something like it.

It was fucking hell.

Alexander Gorny · Xpanner · Original post ↗

Autonomous Construction

If you were supposed to sell shovels during the gold rush, then during the AI rush, you have to sell shovels again. Well, almost. The Korean startup Xpanner deals with excavators, bulldozers, and similar heavy machinery. It takes a completely ordinary machine, retrofits it with its own control unit, and it becomes autonomous: it drives itself, drives piles itself, digs trenches itself, and does whatever else is expected on a construction site.

Monetization is subscription-based, and the unit of measurement is not the physical device, but its capabilities—essentially, apps. If you need to drive those very piles this month, you pay for pile driving; next month, you activate other options. Apparently, for more or less versatile machines, all skills end up being activated eventually, and the startup founders understand this better than their clients.

The project was launched in 2020, and things were apparently going so-so. But then ChatGPT came out. In the US, people went crazy over building data centers, but there's a labor shortage, and Trump is fighting immigrants. That's exactly when things really took off for these autonomous shovels. Xpanner moved to the States, and its revenue is growing exponentially. Recently, it raised its first major funding round in the company's history.

Company website ↗

You can say and think whatever you want, but I know exactly how you feel. Fuck vibe coding. For some of you, fuck life itself. Your jaw clenches with anger. Then comes exhaustion you want to hide in emptiness. Let the emptiness swallow the world. Better that than the constant suffering.

All because of fucking Anthropic. It's playing with your nerves and getting better at it, beating you by an ever larger margin.

I don't know how else to describe the releases this manufacturer of dick cheese keeps treating us to in an already difficult year.

Now I've dropped everything and spent the whole day wondering whether the task will switch to Opus. I sit in one session holding down Ultracode's safety catch, barely breathing, wasting time and the allowances on every other subscription because I can't touch the mouse.

Throwing away an entire prompt pack that's been running for four hours and reverting every change because Fable thought brilliantly but somehow ignored one repo parameter. Without that parameter, all the prompts at best stop Codex in its tracks, and at worst destroy everything that worked. Good thing I noticed in time.

At the end of a day like that, I wonder how this is any better than sewing bucket hats in Bali.

Reliability is a much more important part of quality than peak reasoning. That doesn't mean peak reasoning isn't needed. Just not at the expense of reliability.

Two possibilities: either Anthropic fixes it after seeing people leave, or there's something I don't understand and never will because my brain has too few wrinkles.

One thing will definitely keep happening: the market will shout and mutter about how Anthropic beat everyone, even if they never press another button. Ninety-nine percent of the information out there now comes from people who claim OpenClaw on a Mac mini automated their whole business by itself. There's nothing to be done about that anymore.

I initially took Fable for the usual fucking shit from Anthropic. But before getting into the bath, I tried giving it a task that required thinking broadly. Forty minutes later, I was shocked by what I saw.

Fable's analysis was worth roughly ten GPT PRO analyses. It dug so broadly through the facts in my repo of many thousands of files and connected them to each other.

I have no idea how it codes yet. But I can say no LLM has ever looked this broadly at my repo. The diagnoses are astonishingly accurate. I know all the problems very well, so I can judge exactly what it found and how.

I've spent weeks fighting with GPT to get one enormous infrastructure situation into perfect shape. It's been hard because GPT can't shine its light over such a broad area. I have to define every piece with a contract GPT mustn't cross. Otherwise, it fixes one place while breaking three others.

GPT 5.5 nevertheless sees the situation much better than Opus 4.8. The leap from 4.8 to Fable in this respect really deserves a new name. It's clearly a different class of model.

Still, I don't trust Anthropic. For the last 365 days, every model they've released has gotten ten times dumber by the end of the next day and somehow never gotten smarter again. We'll see.

The limits are fine, by the way. Previously, I paid Anthropic two hundred dollars just for Claude Design and didn't even use five percent of Opus's weekly allowance. Any allowance on even one useful, WORKING model from them is fine.

Maybe the subscription will finally start being useful beyond design. Even if it codes badly, being able to scan a repo like this is already nice. Design plus a search engine and planner.

Although the next GPT will come out soon. Then we'll see the real state of the market for the next month or so.

I asked the dentist who's treating my teeth: “Guys, you're using a website from 1998, calling clients manually, and so on. Want me to build you a fucking great website for free and automate customer interactions? Just because you exist.”

“Uhhh, what the fuck for?” was the answer.

Go make money in that market.

You have to help people who visibly want help. If you first have to convince someone they need help in order to help them, it won't fly.

How often do you think I persuade people to buy Salesbomb? Fewer than ten times in the last seven years. In the overall picture, that conversion is negligible.

The buyers are people who want leads, consciously and strongly enough to pay what they cost. Even then, conversion is seven percent at best.

People don't understand that even when a hundred people come in who desperately need it, only seven to ten will actually pay, at best.

I know plenty of guys who built Atlassian or Salesforce extensions for themselves, and nobody else started using them in ten years.

Alexander Gorny · Ethos · Original post ↗

AI-Powered Matching

An expert marketplace is a favorite startup idea. Company A wants to learn everything about Market B. It can spend six months and a million dollars on experiments, or it can just ask someone who has already worked there. And pay, say, 1,000 bucks for such a conversation. That's a 100,000% return on investment—no joke. The platform that finds the right consultant and processes the payments will take, let's assume, half. The expert gets to keep 500—not bad for an hour of chitchat.

There are plenty of companies worldwide implementing this mechanic with all kinds of variations. The American startup Ethos decided to be the first to add AI to the mix. Unfortunately, they haven't yet figured out how to slip Claude into the conversation instead of actual humans at a rate of $1,000/hour, but they have already entrusted it with selecting the experts.

A voice bot onboards the expert onto the platform. It asks questions about their experience and specialization just like a real person, but unlike a human, it records all the details into the service's knowledge base instead of just a few sentences. Then, when a client submits a request, the algorithm finds a truly suitable expert rather than just a random keyword match. At least, that's the story Ethos is telling.

They also claim to be "on track to $100 million in revenue." You can only believe this if you consider the very beginning of a journey to already be "on track." In reality, the startup is still at a fairly early stage; it raised its first serious round only a month ago.

Company website ↗

I write my own texts.

For good email copy in Knowledge, I had to put seven models in a row, each performing its own function. One cuts out all the AI slop, comes up with synonyms, and picks the most suitable, least embarrassing ones, and so on.

You can split semantics into layers. Logic too. Each layer gets a model. Then one at the end checks how every layer performed and rates how much the whole thing resembles normal writing on a ten-point scale.

That's what it takes for now.

I think tonight's spontaneous migration accidentally demonstrated very clearly how error-free Codex is these days.

I remembered that the old SLSBMB-stats, which was still an analytics platform on steroids rather than Salesbomb automation, was still on Railway and Supabase.

The Pussymoneyweed bot that runs things in this channel and the discussion group, hosts chess and the stream payment mechanism, and handles some other processes was still on Railway and Supabase too.

As of yesterday evening, that meant six servers: one Hetzner at about forty-five dollars for my sender and everything related to it; two load-balanced Hetzners for the mailbox and warm-up farm, also forty-five dollars each; a five-dollar Hetzner for Ultima Online; and two fucking Railway-and-Supabase setups that had somehow started costing about a hundred dollars combined, for no obvious reason.

I've noticed that services selling simplicity for pennies to beginners tend to quietly multiply their prices over short periods for fucking nothing.

Pricing depends on loads of things, including the load my neighbors put on the server. But you can't see the neighbors. They can just keep raising the price forever.

I caught sight of another Railway and Supabase bill and asked Codex: “Bro, what do you think we should do with these repos?”

Codex took a quick look: “Bro, are you kidding? A $4.99 Hetzner wouldn't even reach five percent load with both projects on it. But you don't need to buy another. I'll just move them onto one of the fleet servers. Those servers won't even notice these two projects. They wouldn't notice ten.”

“Is moving difficult?”

“Give me the go-ahead and go to sleep. You'll wake up tomorrow, and everything will have moved.”

I gave the go-ahead and went to sleep.

In the morning, I thought we'd have things to sort out. But no. It moved the projects, deleted the Railway accounts, and couldn't delete Supabase because they won't allow it before the final invoice is paid, but prepared the downgrade to free.

Not a single error found. Two fairly different projects moved onto one server, now shared with a very important project, without one error.

The whole thing took just one hour and five minutes.

May 2026

Anthropic is tormenting me, and I'll keep paying anyway. There's something about that resembling the formula of life itself. At least, it's been like that for as long as I remember.

Opus 4.8 really seems to have gotten smarter and hasn't deteriorated yet. Or not on my tasks. I spent all day using it only on the Ultima Online server. Forty percent of the weekly allowance, and all good. I actually had to work at burning those forty percent while running five sessions in parallel. That's much better than before.

But here's what fucking drives me up the wall about this miserable, piss-soaked fucking shitshow.

The Codex people figured out not to let the model bother me with fucking questions unless I press the Plan Mode button. The Claude Code people didn't.

Now I have to sit there working for this fucking thing instead of the other way around.

Whatever task I set, I'd better not leave the computer. Thirty seconds later, it'll ask its fucking questions.

Questions like: “Do you think I should paint it a little more blue or a little less blue?”

“Or perhaps medium blue.”

“Or I won't paint it. You can paint it yourself, and then we can do something else fun.”

Alexander Gorny · Findd · Original post ↗

On the Way to Going Fully Digital

Monitoring needs to be convenient. When someone has to sign a paper logbook every half hour to prove they cleaned the restroom, it's a drag and nobody likes it. But if a mobile app automatically cross-references geolocation and accelerometer data—now that's impressive, and at the very least, the Chief Digital Officer will definitely be thrilled.

The American startup Findd isn't the final destination yet, but it's a step in exactly that direction. The startup builds automated timesheets for non-desk jobs—janitors, construction workers, security guards, and home health nurses.

An employee arrives at the site and points their phone at their face. The system matches it against biometrics, logs the GPS coordinates, and clocks them in. From there, it tracks the end of the shift, calculates hours based on the relevant union rules, checks if their medical clearance is up to date, and neatly feeds all of this into the payroll system.

Interestingly, a major part of the startup's positioning is its technical sophistication. Perhaps it's implicitly telling potential clients: 'you can't just vibe-code this in two prompts.' Their marketing materials emphasize the complexity of facial recognition and features like offline app functionality—after all, there might not be cell service on a construction site, but you still need to clock in.

In a recent round, Findd raised $21 million—a clear success for a company where the word 'AI' appears on their homepage exactly once, and even then, only in the footer links.

Company website ↗

FINALLY delivered fucking CHECKMATE WITH A PAWN playing both colors in actual games.) It only took three little years and 5,109 Chess.com games :)))

And it was also my record-accuracy game in one-minute bullet =)

I remember deciding to grind chess until I'd checkmated with a pawn as both White and Black, without hints or grinding theory.

But now I've decided to keep grinding chess simply because I like playing it =))

By the way, I don't think I'd have kept my sanity through this vibe-coding hell since January without chess.

If you just stare at the terminal, constantly think in several streams about tasks already running and tasks you need to set, and watch the scrolling progress of each session, your brain gives out. So I work in raids. Open the terminals, check where it's time to add something, force myself to do it, then switch back to chess. That way you can sit there vibe coding for fourteen hours without going insane.)

Something like Dota is too distracting. It's impossible to concentrate on videos or movies. Chess is good because you can control how long it lasts. Losing isn't a big deal if you're not Carlsen. MMR truly is just a number, and so on.

For the first time in my life, I'm seeing mass suffering over 5.5 while nothing has changed for me. What's going on?

It simply does everything I ask on the first attempt. I can sincerely say that every time it doesn't do what I need, it's because I was too lazy to formulate the task properly and got fairly punished for it.

The only negative thing I've noticed about 5.5 lately: under no circumstances should you enable Fast. It's started doing a terrible job while burning three to five times more tokens.

A hundred percent of the weekly allowance doesn't last until evening if I run four Fast sessions. Without Fast, it uses no more than fifty percent.

AND FAST ISN'T FUCKING FAST. The strangest part is that Fast takes longer to do my tasks, not less time.

I think inference runs too quickly and makes more mistakes, but some self-control makes it fix them. So it runs around making and fixing errors for an extra two hours, burning tokens like a bastard.

All right, I'm fucking sick of it. Time to show what I've built and get this unbearable weight of existence off my shoulders. On Sunday or Monday, I'll stream for who-knows-how-many hours to show and explain the technology I've been building all these months.

There's still work to do. But tonight it hit me that there will absolutely always be work to do. Every week I tell myself I'll finish by the end of the week. Instead of celebrating, I invent another two weeks of backlog.

My system has been the best thing on the market for ages, hands down. I have zero doubt that nobody has anything even close.

If someone starts showing off and says they do, they're bullshitting, because they won't be able to show it. I will, on stream.

Today someone in the chat asked: “Anton, are you actually coding, or have you gone crazy and started imagining things?”

Meanwhile, Anton is literally generating tons of leads for real clients with this system. You've got some fucking nerve baiting me. Although it worked. The world belongs to the shameless.

I have a superpower because I wrote my first line of code in September. I don't give a fuck whether code is complex, simple, crooked, or straight. Its appearance doesn't affect me, because I basically can't tell one from another.

Rather than refactor for some incomprehensible reason, I'd run a health check: “How's the infrastructure load? Everything okay?” If yes, everything's wonderful.

If not, both Codex and Claude Code miraculously learn to refactor very well when they have a task like “optimize the server load.” I know that for sure. I've encountered it more than once on Hetzner, especially the $4.99-a-month one.

You set the task: “Please make sure we aren't wasting any capacity. That probably means rebuilding some shit. Rewrite the whole thing if you have to, I don't care. Just make it handle ten times more on ten times less hardware, and please don't give me a fucking headache in the process.”

That prompt looks a bit fucked up at first glance. But the output will be exactly what I asked for. Presumably through whatever refactoring, rewiring, or other fucking thing is needed to fulfill that wish.

I read Darcy Rezac's book. He explains that most people who read the frog-princess fairy tale as children noticed only the romantic part, missing the remarkable mechanics of what happened. The prince kissed every frog in the swamp to find his princess.

The longer you think about that, the more interesting it gets.

I always give people the benefit of the doubt. That doesn't mean assuming someone is amazing. It means assuming the person isn't bad. The business isn't bad. The product isn't useless. And so on.

What happens next will show who the people really are, what the product is, and the rest. How else would you do it?

If someone deceives me, it depends on how they do it and the whole picture. Sometimes you can simply forget the person. Sometimes you can practice your wit and find a new way to insult their fat mother. All sorts of things happen.

But I really can't imagine why anyone would work with people who don't intend to work fairly. One side puts in X; the other should put in its equivalent of X. One side keeps its word; the other should too.

If someone thinks you're shit, continuing to deal with them means agreeing: yes, I'm shit. That's playing with dangerous stuff.

A person owns nothing but themselves. Don't throw away the one real, inseparable treasure you have, because then nothing remains.

The one thing nobody can take from you is what you think. The trouble is, many people give it away themselves. And exchange it for such insignificant things.

Like indigenous people handing over their treasures, wives, and children in exchange for unfamiliar shiny stones and bags of spices.

Don't be that naive. The pay is fucking terrible.

Codex wrote me some kind of Telegram bridge. It works like clockwork.

With every message I send, it opens a new CLI session but feeds it into the bridge's parent session using that session's ID. That way Codex remembers the previous several days of my requests and works perfectly on the computer, just as if I were setting tasks in the Codex app.

The only thing is that I need to say which of my repositories the project is in. It would figure it out itself, but spend an extra minute doing so.

It even emits its thinking process like the CLI. That wasn't available out of the box. It had to write some extra code to emit it.

Vibe coding is a skill, a process that continues over time.

When we have AGI, you'll simply ask for what you need, not even HOW, just what, and you'll get it. No skill will be needed. If something requires no skill, it no longer has any value; it loses its meaning in human terms.

For example, you drive a car: that's driving. She's a better driver than he is, and he's a better driver than that person. But nobody treats riding a bus as a passenger as a skill called “bus-riding.” You can't say, “He rides the bus better than she does.” Well, you can, but the guys in white coats won't approve.

Right now vibe coding involves what, how, in what way, by what means, and so on. Lots of questions equals a skill. One vibe coder can be hundreds of times more in demand than another. There's a reason to talk about it.

Alexander Gorny · Descript · Original post ↗

The Heavy March of Progress

The American startup Descript automates video editing. You upload a podcast video, write in plain human language—'remove pauses and clean up filler words'—and it actually does it. If you write something more complex, the startup promises to handle that too, but I haven't tried that.

Descript did a pretty good job with the initial processing. Out of 15 minutes of my speech, it cut out a minute and a half, without cutting anything it shouldn't have (or at least I didn't notice it did), making it much more pleasant to listen to. In a couple of places, I didn't like the cut, and I definitely would have called out a human editor for that—money was paid, but the work was sloppy. Here, I just let it slide. Maybe I should have said something; it might have been fixed on the second run. I don't know.

The cheapest paid tier costs 24 dollars a month. The limits are enough to process 10 hours of recording—you can't get much cheaper than that. And, you know, progress is scary. There used to be an entire profession. I used to look for people, pay them. I waited for them, by the way. And now, click-click—and everything is ready for 24 dollars. What's even scarier is that Claude Code did the exact same thing with another video. Without a separate startup and separate money. Just out of the box. Almost in a single prompt. Essentially for free, as part of a general subscription. Where are we heading with this? I have no idea.

Descript's last funding round was in 2023, when the company was valued at around half a billion dollars.

Company website ↗

Alexander Gorny · Omni · Original post ↗

Replacing Humans

For about 15 years now, we've been promised that with the next new tool, analysts won't be needed: "the manager just asks the computer, and our software draws beautiful charts." Previously, this was promised by BI dashboards—meaning specialized report-generation software—and since 2023, by AI add-ons built around BI dashboards.

The American startup Omni continues this fine tradition and promises the same. Its hypothesis is that a human is still needed to provide context. When a real, live analyst receives a request to "calculate active clients from Chile," they know that "active" in this company refers to those who made a purchase in the last week, and they will answer based on exactly that. Meanwhile, some hypothetical ChatGPT might decide today that it means "those who visited the website in the last month," and tomorrow "those with an active subscription." And I haven't even nitpicked "from Chile" yet—which is also highly ambiguous. You could use, for instance, the delivery address, phone number, country of passport issuance, registration IP, or last login IP. In short, AI will certainly calculate something, but the numbers will be misleading.

Omni, accordingly, builds an internal knowledge base within the company containing these kinds of definitions and accepted corporate rules, and then writes SQL queries based on that rather than just pulling them out of thin air. The manager asks a question in natural language and gets charts that aren't just formally correct, but actually useful to them. Analysts are not needed. At least, that's what the startup promises.

In a recent round, it raised 120 million dollars in investment, and its valuation reached one and a half billion.

Company website ↗

For a long time now, at the end of every run my model has had to check whether everything reached production, whether main is clean, and whether every process launched in that session has been killed: Playwright, Chrome, scripts, and anything else that consumes resources and won't stop itself.

Since I asked it to add that everywhere necessary, I don't even know exactly where, there haven't been any more problems.

I sometimes spend five hours on a plan, or ten. Depends on what's being built in that plan, of course. But a good plan usually costs no more than 80% of one weekly token limit, and there's not much to clean up afterward.

Doing it manually would therefore have turned out just as badly as through a loop.

The problem with /goal is that it does whatever it wants with anything that isn't defined. And what it wants is to minimize the time spent. Unless how everything should work is described in the tiniest detail, it might just make the whole thing out of papier-mâché.

I think /goal simply isn't finished yet. It's still an experiment. It's missing some kind of engine in front of Goal that formats the task properly for Goal. A human can't handle that role.

It needs some kind of Ultra Planning Mode. Which is why goal hasn't appeared in the app.

Alexander Gorny · Neato · Original post ↗

Gen Z Invented the Middleman

Once upon a time, life was good for brands. Amazon would buy the product, sell it to the end customer, and the manufacturer didn't have to worry about a thing, just count their money. Then the store turned into a marketplace, and life went downhill.

Either you are being resold by hundreds of nameless sellers—where everyone's descriptions and photos are terrible, counterfeits sit right next to the original, and the customer feels like they're at a chaotic flea market rather than a flagship store. Or, you have to deal with Amazon yourself: manage ads and product listings, fight off competitors—a type of business that many brands are simply not cut out for.

The American startup Neato offers a middle ground, reviving the classic distributor model in the era of marketplaces. It simply buys inventory wholesale and then handles the promotion on Amazon and other platforms itself, taking care of pricing, discounts, advertising costs, logistics, and fees. It also tries to chase away independent sellers as soon as they violate any rules. In return, it gets the traditional distributor margin—the difference between wholesale and retail.

The startup focuses on everyday consumer goods: pet food, kitchenware, groceries, cosmetics, dietary supplements, and the like. The product must already exist and have some brand recognition, but either there isn't enough money yet for an in-house "Amazon department," or the owner trusts a contractor more than their own employees.

In a recent round, Neato raised $25 million in investment—which isn't much for its business model, considering it needs to buy inventory. Then again, maybe they just aren't telling the press about their purchasing loans.

Company website ↗

Just remember that “a thought once uttered is a lie.” To understand it better, adapt it to the modern internet: “a thought once uttered is a point of view.” Or an opinion.

Someone who believes you can and should manage only slaves probably sincerely believes it. That doesn't mean it's the only way.

That's a problem with our world: there are opinions and there are formulas, and outside the exact sciences it's very difficult to tell them apart.

An opinion works well only in the life of a compatible cohort. Say someone knows how to DOMINATE. People who want to SUBMIT need that person. But these are cohorts. He'll insist nothing else works, when nothing else works FOR HIM, because he may not be strong outside his cohort.

There are also logical relationships that work regardless of cohort. For example, average email conversion across the internet will be higher or lower depending on the approach to content. That approach is an extremely complex formula with an enormous number of variables.

But no science has established the formulas or variables as fundamentals.

Tomorrow, for example, I'll talk about what I know better than anything in life, better than the back of my hand: writing or reasoning in online sales. It won't be an opinion. It'll be hard-earned experience, measurements, real numbers.

For most people, though, it'll remain merely an opinion, a point of view.

Every day we encounter a fucking enormous amount of information that's either opinion or physics. Distinguishing the two every time would be a tremendous gift of reason, but nobody can do it across the whole flow. The more developed that ability is, though, the stronger the person is in relation to the market, society, and so on.

Since I'm chewing an M&M while staring at the monitor, with neither a mouse nor a penis in my hands, I can write up an interesting observation. It's really your observation. You see it every day.

People constantly argue that AI won't be introduced into serious establishments in Wyoming or the Russian city of Syktyvkar because AI isn't accountable. Only a person can be responsible, therefore a person is an indispensable part of any system.

A post I read two minutes ago gives this example: imagine a server room. At some point AI goes crazy and deletes your whole server and all the data you've carefully collected over ten years of serving clients. AI isn't a person and won't even bear responsibility. Therefore there will always be a person there.

Do the people passing around this tired refrain like a Soviet-era speech, ever even try to visualize the situation?

Suppose the person who accidentally deleted everything does bear responsibility. What responsibility? Will they fire him? Fire him and scold him first? Fire him, scold him, AND TAKE AWAY HIS ACCUMULATED BONUS?

Maybe flog him? Beat him up behind the garages? What else will soothe the system that's just lost its server and assuage its longing for the lost data?

Or am I missing something about life? Will that shaggy-haired fucker in the server room bear such immense responsibility that the shareholders are just waiting for him to delete their data and pay them two dollars for every dollar of lost shares as punishment?

There's the exit strategy. Wait for our employee to screw up and bring us a billion as punishment!

But he doesn't have a billion. Or a million. He has fuck-all. Even if he has something, he won't give it to you. He might wave his dick goodbye, then find another job.

However much I read about AI versus responsibility, I've never understood what people are talking about. Maybe you can explain it.

Please tell me what difference it makes whether some fucker deletes my data or GPT does. I really want to understand, but I can't, because I'm an uneducated idiot.

This thing sucked a lot of blood out of me. But now, finally, I'll just drop in a client's domain, a call recording in any file type as long as Whisper can make out our mumbling, and any random documents, all dragged into one junk-drawer folder.

A dedicated Codex session will spend two days researching the market. It'll produce the most thoroughly researched set of ideas imaginable, deepening the weighting of hypotheses and discovered evidence every hour, mapping job titles and subtle relationships between company sizes, roles, and markets. It'll analyze every competitor and their competitors, direct and indirect.

All that becomes an endless queue of campaigns, tested against one another and promoted to the next tier of development with more contacts, or disqualified after repeatedly losing to other hypotheses on equal sending volumes and their telemetry.

Who would have thought this would ever become possible?

McKinsey charges millions of dollars for reports it takes months to produce, worse than what a system like this will do practically for free in a couple of days.

That's not a figure of speech. I've literally held McKinsey reports that someone paid a million dollars for. Actually, they paid two million. They were such a fucking disaster I remember opening a bottle of whiskey.

I'd have opened it anyway, though. What the fuck else is it for? Standing on a shelf?

April 2026

I've got Nano under control, by the way. It can do MUCH more than I could have dreamed, and it costs fuck-all. Twenty-five thousand requests a day comes to around eight dollars a week if you use it optimally. That's basically free, and there's enough intelligence there.

Mini is very convenient in that respect too.

The problem starts when you want PRO intelligence or extended thinking. But in my tasks, amusingly, extended thinking turned out to be necessary only where I couldn't replace it with my own extended thinking.

I'm capable of sitting down once and writing a long, detailed explanation for a model about my domain expertise. Once that's there, thinking is no longer needed. With an extremely long, detailed instruction, Nano handles the task with 99.9% accuracy.

To get to 100, it's enough to require not one green verification flag against the result in the database, but two or three. Nano reaches 100% in three passes for sure.

And that's still tens of thousands of requests per dollar.

I'll calm down when this flow works exactly like this, without errors.

The client pays. We catch the payment, create a workspace, automatically register the infrastructure and start warming it up. We launch deep research into everything related to the client's business, since we know their domain. We launch deep research into their competitors' businesses, their clients, and their clients' clients. We merge all of this into a specific database for the client. This is where we are now.

We have a call with the client and talk. I simply drop the resulting call recording into the client's workspace as a file. The database gets enriched. Once all the checkboxes are filled, a cloud of hypotheses appears. A crawler collects contacts for each hypothesis and runs them through a sequence of verification services for 100% agreement. AI writes four versions of the email and four versions of the A/B tests through step eight.

AI creates campaigns for the first 50 verticals, fills them with content and contacts, and starts sending emails. All incoming emails are forwarded to the client and displayed on every dashboard with all the necessary information about the person, company, and industry.

That's V1. V2 should automatically correspond with leads, getting them to a call and then following up afterward. V3 should automatically conduct the calls.

I think I've just reproduced the current difference in intelligence, tool use, and so on between GPT 5.5 and Opus 4.7 with a very specific example.

I have a new Hetzner Robot server where a system consisting of seventy-eight tasks is being built. I decided to do it entirely through Claude Code /Loop to see how that process builds a whole system.

Loop did a lot. Not all of it correctly, but we can live with that. We'll fix it.

About sixty percent of the way through, though, Loop started stopping because it couldn't do certain things independently after three attempts.

The latest obstacle: properly set up and bootstrap the new Stalwart 0.16, then connect all the systems already running and configured on that Hetzner server.

Claude first suggested taking the task over from Loop, supposedly because Loop had something wrong. But Loop is the same Claude 4.7 Max. Predictably, the parent orchestration session couldn't fucking do it either.

Then it said I had to do it by clicking through everything myself. The machine was helpless against all those menus, tokens, exports, and other unfriendly Stalwart bullshit.

I asked Claude to pour its suffering into a separate file for Codex. Then I asked Codex to read it and help the poor thing.

GPT 5.5 xHigh managed in thirty-two minutes, including tests of every part of the system. It also fixed about five configuration bugs Claude had introduced earlier.

Claude looked at that and said it hadn't expected another model to succeed. It was tearing its hair out and wondering whether to rebuild Loop so it delegates every difficult task to GPT 5.5.

Vibe-coded it. Bro, a CRM is already one prompt this year. Including the API.

Don't picture what I'm doing as a collection of tools. I'm making ONE tool with a million features.

Here's how I go about it. First there are tons of buttons and a visual mess of shit. That's okay for working things out. Then I polish it toward a state where, ideally, there are no buttons at all.

You don't immediately know what to automate completely. You have to live through it. But once I understand where the different parts of the process meet, I can almost always reach an understanding of how it should eventually work by itself.

And it gets pushed out of the UI. For example, I wiped every invoicing process out of the UI. It just works now.

It'll be the same with everything.

Yes, I do that. Codex usually says there's one idea that could be useful, and we absorb it.

The other day someone sent me the Y Combinator CEO's full setup: plugins, skills, MD files, and so on. Claude found four useful things in all that stuff. Codex found three. Seven in total.

They rejected the other 23,426,937,462 things, saying they were crap and ours was already at least as good.

It's simple prompting: ask them to compare it with our world and work out what we're missing and where ours is better. It's also a good time to clean up and ask them to remove principles that have become outdated or contradict each other.

Alexander Gorny · ALSO · Original post ↗

Squat delivery robots that look like old Soviet canister vacuum cleaners are actively cruising around the world's megacities. They have already become an almost established standard for autonomous couriers.

And it's a very strange standard. The rover is small, so it can't carry many orders. The rover is slow, so the food gets cold. The rover doesn't have much range, meaning it's only good for short routes. Ultimately, the rover is just parodying a human pedestrian, but real-life delivery people don't work on foot. They use bicycles.

The American startup ALSO makes several products, and today we're looking at their autonomous electric bicycle. The vehicle itself is large, so it can hold a lot of cargo; its wheels are big, so it rides fast; and its battery is powerful, allowing it to handle long-distance deliveries. To my taste, the device looks more like a cart, but the company prefers the term 'bicycle'—they want to be allowed to ride it on bike lanes and, apparently, even on sidewalks. This is safer for the delivery vehicle, and the autopilot is easier to program when you don't have to worry about the high speeds of surrounding traffic. Regular people probably won't appreciate sharing space with it, of course. Then again, they don't like actual human couriers on bikes either, so it won't get radically worse.

It seems ALSO hasn't delivered its first order yet, but the product is far from just a raw concept. A contract with DoorDash, the largest American food delivery service, has already been signed, and the startup raised a whopping $200 million in a recent investment round. Part of that funding will go specifically toward this bicycle courier.

Company website ↗

Nothing has changed. Restrictions only work properly at the AGENTS.md level. And you really shouldn't overdo them, to put it mildly.

In prompting, it's better not to get into the model's business with “how,” and to focus on “what” and “why.” It might not be able to do something the way you want at all, but could easily implement it another way available to it. You've already ruled that out by constraining it too strictly.

I'm sure coders fall into this trap constantly.

I don't quite understand which tasks you're talking about or why track them. If something occurs to me, I ask for it to be implemented immediately. Then it can't get lost.

Sometimes it works straight away. Sometimes it works badly or doesn't work. But once it's in the repo, its turn, someone's attention, or a refactor will reach it sooner or later.

What's the point of a task you aren't ready to ask the model to implement right now? If you can't describe it well enough to give it to the model, maybe it doesn't exist and you just imagined it.

As I recall, there was research saying a person's stress level is directly proportional to the number of unfinished tasks their brain knows about.

Alexander Gorny · FirmPilot · Original post ↗

The Difference Between a Rat and a Hamster

A full-service marketing agency offers to build a website for a client, write SEO articles, launch search ads, manage social media, and monitor reviews on maps and directories. This is a dull, low-margin business. Neither the press nor investors will ever care about it.

The American startup FirmPilot offers to build a website for a client, write SEO articles, launch search ads, manage social media, and monitor reviews on maps and directories. This is a venture-backed startup that recently raised 22 million dollars in funding, and even I am writing about it.

The main question, of course, is what the difference between them actually is. FirmPilot claims that AI-AI-AI in all their processes has allowed them to deliver better quality at a lower price. They focus on serving law firms and charge them 4,000 to 9,000 dollars a month—quite cheap by market standards. A skeptic would argue that anyone can dump prices using venture capital, and everyone is already using AI in pretty much the same way, making it a market norm rather than a differentiator.

However, a compromise can be found between these two extremes. New tools will redistribute the market; some won't implement them in time, while others will do it better. And, obviously, with funding, there is a better chance of ending up on the right side of history and capturing a large enough market share to recoup the investment. Maybe it will all actually work out.

Company website ↗

What I've been rewriting from scratch for the fifth time since the 23rd is an example of something difficult to develop.

The tests alone are already 112,000 lines of code. TESTS.

Ales, I've realized the difficulty starts not with complex solutions, but with many connections between solutions that may individually be fairly simple or medium-complexity. Every fucking thing in my system interacts with dozens of other elements through pretty tangled logic.

Say each element is a thousand lines of code. Simple arithmetic suggests twenty thousand in total. In practice, it easily becomes two hundred thousand.

Every class of solution also needs its own approach to security, resilience, self-healing mechanisms, end-to-end testing, reactions to unexpected test results, and so on. All of it is code, code, code.

Back in 2020, COVID happened, our whole civilization collectively lost its shit, and I started wondering what kind of work chair best suited the situation.

Before that, I didn't have a chair. I was used to working on benches, beanbags, hammocks, the floor, and, of course, all kinds of sofas.

So I asked Facebook about chairs, which I knew absolutely fuck-all about. GPT hadn't been invented yet.

The only thing I knew for sure was that I didn't want the embarrassing Herman Miller Aeron. I'd had one at Yandex, and I wanted nothing I'd had at Yandex. Sitting in that gray monstrosity felt like returning to the open-plan office, surrounded by show-off nerds trying to look smart while doing nothing. They kept distracting me because being next to someone actually working made them uncomfortable.

People suggested various things. I bought various things. For some reason, the favorite was the IKEA Markus. The worst fucking chair I, or anyone, has ever owned. I could have sat longer on an ironing board, which didn't look much different.

Then I tried chairs whose names I don't remember. Ergonomic ones, gaming ones, even OFFICE ones.

Of course, I eventually sat down in a Herman Miller. It's the only chair I don't feel in the second, eighth, or sixteenth hour of work.

That experience let me define the perfect chair: you use it every day without remembering it's there.

You can buy a headrest for the Herman Miller. A place to put your big fucking head, basically. Ads targeted me from the moment I ordered my Aeron online, but I stubbornly ignored them for years.

Eventually, I gave in and bought one. I'd gotten an excellent firm footstool for under the desk and imagined turning my Aeron into a sofa for watching strategically important YouTube videos. Lean back, which you can't do without somewhere to put your head.

Naturally, the headrest doesn't just look like shit. It is shit. It's very comfortable, like the Aeron. That's where its advantages end.

Herman Miller's engineers managed to outdo themselves and create “Very Human Design,” like that Chinese TikToker's absurd inventions.

The whole headrest attaches with one bolt tightened by an Allen wrench. That bolt has to clamp it firmly to the chair, preferably until something crunches. The innovative plastic is very strong but still flexible. Stretch, and an insufficiently tightened headrest falls off. At night, it sounds like a fucking bomb.

You'd think: where's the problem? Tighten it and shut up.

But there's no wrench that fits that hex socket except the one supplied with it.

I didn't think that was even possible. It looks like an ordinary hex socket. Stick any hex-shaped thing in there and off you go.

Fuck no.

Naturally, the original wrench, the one that actually worked, gets lost within the first few months. But I'm a man. I have a shitload of tools.

So many that I can easily surprise any builder who comes over to fix something. Whatever he's missing, I generally have it.

A detector for wires in the wall? No fucking problem. A voltage tester? Here, the best one. A hacksaw? Take your pick. I have two.

Sanders, dehumidifiers, six kinds of silicone, every type of paint used on every surface in the house. Want to sand something? Eighty grit or six hundred? Everything in between? Of course I have it. Want a power sander too? Here.

I can drill through any surface to any depth. When a guy broke his ridiculous drill bit for the internet cable and said we'd have to wait until tomorrow because it was his last one, I handed him one from my supplies. He said that in over ten years installing internet connections, he'd never encountered someone with that fucking drill bit at home.

Of course, I have a billion different Allen wrenches. From furniture, DIY kits, all sorts of places. More than fifty altogether.

Not one fits this fucking thing. They either won't go in or spin freely.

When I told the builders fixing the bathroom, they didn't believe me. Said that couldn't happen. I agreed. But there was the miracle, right in front of us.

They tried every Allen wrench they had. Even brought their entire toolkit over from another job. Nothing fucking fit.

Herman Miller can really pull it off.

With all those tools, of course I have a file. I could turn a wrench that doesn't fit into one that does.

But here's what I thought.

When I get my first humanoid robot, just as I once got a robot vacuum, this one use case alone—digitally catalog every tiny object in my home and bring it to me on request—is already a fucking miracle worth buying the thing for. Before considering anything else it can do.

The right wrench is lying around somewhere. I'll just never summon the enthusiasm to turn the whole fucking apartment upside down looking for it. It isn't worth it. Easier to say fuck it.

But soon there'll be a robot that knows how many washers of each size I have, how many milliliters of mineral spirits remain, whether there's a nail of the right color for the material. If something's missing, it'll suggest a substitute. Or take the fucking file and make me the right-sized Allen wrench.

It's already like that on the computer. This year, I haven't searched for anything myself or gone digging through settings. I do it all through Codex. I remember with horror spending a fucking hour watching YouTube videos just to configure something properly.

I installed the Atlas browser as soon as it came out and can't understand how I lived in ordinary Chrome. I still have Chrome. Chrome DevTools opens it. But in six months, I've forgotten how to use it. You type something, and there's fucking Google, and links, and you have to click them.

I can't wait for these things to reach the physical world beyond YouTube, laboratories, Amazon warehouses, and Hyundai factories. Into our houses and apartments.

It's become clear what all those people at all those Instantly-like companies have spent so many years doing.

There are more microscopic cracks in this technology segment's ass than stars in our corner of the universe. Any one thing rubs against so many other interdependent features that every little detail means changes across thirty to seventy repo files.

Right now I'm refining the Qwen email sorter in the unified inbox so there's no situation where an email of any type arrives and Qwen doesn't know where to put it.

That touches a fucking mountain of logic, although it sounds like you just prompt Qwen and tell it which table to sort into. Do only that, and nothing works. It has to recheck the entire chain of activity, workers have to back it up, and another worker has to periodically recheck the sorting, backfilling weeks, and so on.

Yes, you need to use PRO in a fresh session.

My personal, subconscious hunch is that PRO is forced to be more meticulous about every detail entering its prompt. Models receive the entire previous history along with the prompt. PRO specifically gets worse at filtering what's new and what's merely background once its context exceeds a certain size.

That happens very quickly in its case, because it takes huge inputs and produces huge outputs. I think it starts after compaction. For me, compaction happens on the very first request because I feed it whole repos and ask deeply probing questions, expecting a 200-kilobyte MD file of instructions back.

Why I love and hate vibe coding: you think of two tasks, then sit waiting for 246 to finish.

While Claude does them, I keep playing with features. Compared with Claude, Codex basically has no features at all.

I've now set up an always-on terminal session with root access, nailed to a Telegram bot. I also installed Whisper V3 Turbo so the model catches even a sleepy mumble, the breathless voice of a hardcore runner, or the sacred voice of someone whose mouth is occupied by beer.

Then I added the newest, most fashionable, stylish, fucking excellent TTS on top. Surely I'm not going to read what it coded for me.

I know lots of people in the chat have used these features for ages and are used to them. But walking the dog and building a couple of features through “Hey Siri” without touching any device felt like stepping into outer space to me.

Alexander Gorny · Rec Room · Original post ↗

The Evolution of the Metaverse

Roblox is one of the most successful models in the entertainment industry: users make the content themselves, play what others have made, and the platform collects a commission on every transaction. The cost of goods sold is essentially zero; volunteers and partners handle the production.

Naturally, the model is attracting copycats. The American startup Rec Room is building Roblox for virtual reality. You put on a VR headset and enter a cartoonish world with thousands of rooms and mini-games built by other users, where you play, socialize, and build your own things. Mind you, you can also play without a headset—there are versions for phones and regular computers—but VR is the key marketing promise and differentiator from Roblox itself.

The startup's economy is structured like a standard marketplace: the internal currency, tokens, is bought with real money and spent on other players' content. The App Store or Google Play takes 30%, Rec Room gets 30%, and the creator withdraws 40% to their bank account. Platform fees are expensive, indeed.

Naturally, the startup fit perfectly into the metaverse hype. At the end of 2021, its valuation stood at a fantastic $3.5 billion, and the total volume of its funding rounds had almost reached $300 million. With investors' money, Rec Room attracted a whopping 150 million users.

Unfortunately, money can't buy interest. It turns out all those users didn't actually need headsets and VR. The company recently announced its closure—the servers will be shut down on June 1st. Snap picked up part of the team and assets for pennies. Not all startups are successful.

Company website ↗

So far, flying with Claude is going fine. Fine enough that I've already paid for the two-hundred-dollar plan.

People warned that Opus burns through too much too quickly, but I don't see a difference from 5.4.

Problems such as queuing come up and get resolved immediately. You can ask Claude for more ambitious things with tools, because its tooling is noticeably better and it has far more tools.

I told it I wanted it to achieve certain targets. I didn't want to steer it. I wanted to sleep. It came up with nightbench-mode.md, describing a loop for reaching those targets, and wrote a script that starts a new task as soon as the previous one finishes. Every new task begins by reading nightbench.

Very sensible. It does the work, researches, records progress, and moves forward itself.

On one hand, having no prompt queue is unsettling. On the other, what we've made is much better than a prompt queue. I don't think Codex can do that. At least, I never got there “accidentally and intuitively” with Codex.

It's also nice that Claude doesn't give a fuck what I'm doing. Email campaigns don't bother it.

I don't want to draw any big conclusions yet. New things always seem a little warmer, and that feeling is often misleading.

But 5.4 feels more like Sonnet than Opus. Opus seems smarter. And it's definitely faster. In three hours, Opus knocked out as many tasks as 5.4 would have done all night.

Naturally, I use only Opus and only Max, or whatever the highest thinking level is called. I tried Sonnet. It can fuck off. Why use anything but the best? We only live once.

The ideal setup might be paying for GPT PRO and the maximum Claude plan instead of two GPT PRO subscriptions as I did before. With a little Thai Qwen on the side as backup.

Redneck vibe coding is getting more fun.

March 2026

Add the fact that something that takes two weeks today will take one prompt in two months.

Investing in your systems today means investing in system design, not code. Maintaining, strengthening, and improving it will get cheaper as neural networks get smarter.

Here's my own simple example. On the 22nd, I decided to completely rebuild my sender and campaign-running interface. Today is the 29th. Everything's ready; I'm just making improvements now.

But I wouldn't have managed it in seven days without already having the design of the system I needed, which I'd previously spent more than a month on.

When the model nicknamed Potato comes out, I'll rebuild the system again so the code is as good as an LLM can make it. This time I'll do it with one prompt.

If someone wants to recreate my system in their garage, they'll have to spend months figuring out the feature set and learning all the pitfalls the hard way.

Only things I physically can't make, like a MacBook. Everything else gets knocked off as I build it. When its turn comes, I'll make my own Zoom too.

There are plenty of reasons that's better. One example: when you don't own pieces of the ecosystem, cross-integrating them becomes orders of magnitude harder, sometimes impossible. Conversely, when you own every piece, you can integrate everything to infinite depth.

The ultimate goal is essentially what OpenClaw promised, except it was bullshitting, because promising isn't enough.

I want any task that repeats more than once per period to happen automatically. At most, through a casually tossed-off short command in natural language.

If every piece of what you do stores 100% of its data in your own databases, that becomes possible, easy, error-free, uninterrupted, and so on. No need to think about other people's API limits or track changes to them.

The market can drop all the way to zero. Antifragility at maximum.

It can, of course. But I suspect the way files are read depends entirely on how the prompt is formulated.

GPT PRO does a great job of writing a complex prompt. But where GPT PRO loses to a person like me is when I describe the prompt as a full-page confession.

The problem always gets solved much better and much more deeply. Far more tokens get spent on a tiny fraction of the problem.

I think someone sitting down to pour out their soul at length manages to scatter invisible crumbs of context that a formal task brief badly lacks. Even if the formal brief is entirely correct, it misses something that comes bundled with emotional writing and affects how the LLM perceives the scope of the task.

In all these years, I'd never personally touched free models like Llama. I thought whenever people did anything involving reasoning in their services, even something elementary, they were spending money on tokens.

It turned out to be completely straightforward and dead simple.

If you need your service to do something with a touch of intelligence, such as sorting emails into the right folders, as in my sender, it's as simple as asking Codex: “Please wire in the latest Llama so it decides which folder to put every incoming email in. Make sure you describe the folders well so it doesn't have to guess what we mean by soft bounces, automatic replies, or reply-intent types.”

I thought I'd have to configure and study something, but I didn't. It simply plugged in the model, and from the very first test, the model did everything exactly as expected. For free.

My strongest side in life, a self-diagnosed talent or deviation, because it certainly brings no less pain than benefit, is the ability to see a matrix of parameters around whatever's being discussed. It renders so quickly that each new letter changes the entire picture.

It makes it hard for friends, Masha, colleagues, or anyone to talk to me. Talking to me is unpleasant if I let loose without restraining myself with every fucking unimaginable ounce of effort. Because when people say words, they mean something. As those words fill the seconds, without any effort at all, I see not just the matrix of possible meanings but a multitude of branches into the subjects they touch. Countless variations in how each parameter affects each of them, and onward for an unknown number of steps.

That's how a conversation about subject X expands into an enormous set of meanings, from which a completely unexpected, to put it mildly unrelated, subject Y gets pulled out. Without a pause, I tell stories K, J, I about it, chosen from a fucking enormous set of real, embellished, or entirely invented stories. Filtered on the fly through whatever matters in the moment. Stories that can nevertheless contribute to the original purpose of discussing X at that particular moment. The matrix of available parameters is absolutely gargantuan. Even the shape of an imaginary conversation partner's eyelashes matters.

It all happens in an instant. There's no delay. A delay only appears when I have to make a complicated choice inside choices inside choices.

I think everyone's brain can do exactly the same. But not everyone has the kind of access to it in communication and in processing written and spoken language that people like me have. Just as not everyone can hear musical notes or tolerate a surplus of adrenaline.

It's easier to succeed with this affliction because you always have trillions more possible actions appropriate to any situation. But living is much harder, because going out among people is fucking unbelievable slow motion. There are no words for the kind of pain it causes.

And you have feelings for people. The deepest ones will never find a response, because by the time it arrives, you've already changed your mind six hundred and eighty thousand times about where to put the mailbox. And you're completely certain those deepest, most sincere feelings have already gone unanswered.

The thing I hear most often from the people closest to me is that my mood keeps changing radically. But inside my head, I don't feel like I have a mood at all. I'm just analyzing data and reacting to the pattern I'm currently observing. If it changes every millisecond, I can change my response every fucking millisecond. How else?

How do you fucking react when someone puts a cigarette out on your arm? And when chocolate melts in your mouth? That's how I react to absolutely everything coming at me, including internal voices. Only it's happening all the time I'm awake, at a density I can't even describe in words, though I'm trying.

Honestly, I've gone a little fucking crazy with this vibe coding.

It's all fascinating.

But it's killing me.

Killing me like cigarettes do. Or like a good sci-fi book you can't put down when you really need to get up and go to work. Instead, you keep turning pages to find out what the fuck happened to his spacesuit when he finally got hit, and the enemies who jumped around the corner to finish him off found an empty space. There was nowhere left for him to crawl, but he'd disappeared...

What's happening to me right now is completely unnatural. I deliberately developed a huge range of skills, enough to hold my own in front of the Khodorkovskys of this world and actual prime ministers. But none of those skills were developed for what I've been doing these last few months.

You can only get into something like this by accident. And that accident is an important part of the gripping drama in your head that won't let go, keeps pulling you forward. Like wanting to “try climbing over there too” on an abandoned construction site as a kid. Not just because it means something, respect from the boys, a flutter in the girls' hearts, but something for yourself. You're genuinely fucking curious what's over there!!!

But what gives all of it a special charm is another very childhood feeling: knowing it'll pass and never happen again, to you or anyone else.

The vibe coding I excitedly experienced in September 2025 isn't even a pager in the age of the iPhone 17 anymore. It's already a punch card. That feeling of kinship with the handful out of millions who soldered their motherboards to support eight more megabytes of RAM, then went to an eight-person forum meetup in a bar on the outskirts to show off their experiences: you can't mistake it for anything else.

Everything we're doing on 5.4 xhigh today won't make any sense in even three months, let alone three years or thirteen. You won't even be able to explain to the kids of that time what it was or why. The occasion won't come up, just as I never find myself explaining to children how we rewound cassette tapes with a pencil.

The only people who can understand this feeling are those hopelessly afflicted with curiosity, whose symptom is enthusiasm.

It's also important to understand that my approach to LLM development may differ from yours. I definitely spend fewer tokens than many people.

After every prompt, I run two or three investigations that are heavily tool-oriented. Tools barely consume context, which means they don't consume credits. I also have a log-checking step twice between prompts. The model waits ten minutes doing nothing, then checks the logs and fixes problems if there are any.

Without all those intermediate stages, pure coding would probably have burned more than 50–60% over the same period. But that's a pointless way to code: the error rate rises a lot, and cleaning everything up is exhausting.

I've learned not to dismiss a few seconds spent on LLM development. Or minutes. Or even a couple of hours. Those hours can save days in the worst case and still pay for themselves in the best case.

The problem might be in the repo description.

I have a global skill for all of Codex, regardless of the repo, that makes it read README.md and AGENTS.md before starting a task. After working on the repo, it has to think about what we've significantly affected that needs updating in those two files.

Without that, I can imagine Codex having to read through the repo to get its bearings. Mine just reads those two files and knows everything. It also knows they're current because it sees the notes, and it sees the whole prompt, including the system instructions and every active skill. So it knows the updates definitely happen. That affects its behavior.

In general, every little detail affects the model's behavior globally.

But don't just indiscriminately make documentation. I had a phase of stuffing in research, whole batches of API documentation, and so on. You can fill the entire context with that shit and have nothing left to work with.

One fucking excellent README.md is enough. Get PRO to think through its structure, because what we think a model needs and what it actually needs are different things.

When GPT PRO writes me a serious prompt for Codex, I read it and think, “Fuck, there's no way that's enough for it to do what I want.”

But it is.

Every repo of mine has a Governer skill whose job is to check the repo for fucking trash: no temporary workarounds, no research for some task from yesterday, and no files that aren't used in production.

A Codex automation triggers this skill several times a day.

By the way, as you know, Nikita and I got seriously fucking wrecked a couple of days ago. Nikita's a hardcore drinker in the Irish sense of the word. Next time I meet him, I'm wearing a muzzle, so I can still get very drunk without actually dying. Second time drinking with him, second time I don't remember cycling home, even though my memory leaves my body no more than once a year. I'm sure the bike will remember that ride for the rest of its life.

Nothing to brag about here, except one thing: I was certain a proper hard drinking session would break through the nicotine barricade provided by my 25 mg NICORETTE patch.

But after two pints of beer, I was holding up without even trying. After four, I smelled smoke in the air and thought, “Holy fucking shit,” then immediately forgot about it. After the fifth pint, a cigarette crossed my mind, but that fleeting memory instantly drowned in a shot of Macallan.

The sixth pint was already mixed with Macallan, so it's hard to say whether I wanted a cigarette by then.

BUT I DIDN'T SMOKE A SINGLE ONE.

So today is either my fifth or sixth day without smoking. And I'm almost uncomfortable with the thought that, honestly, I don't particularly want to. Normally, you understand, I wanted it considerably more than air =)

This is the longest break from cigarettes in my life. Since I started smoking eighteen years ago, I've never gone more than about 42 hours without one.

If you smoke and feel bad about it, buy the fucking patch and enjoy yourself. By day five, I can already smell things I couldn't smell before, and food has filled up with fucking vivid flavors =)

I'm waiting for the great moment when life takes these newly sharpened senses to McDonald's. I'll bite into that Quarter Pounder with double cheese and cry like a girl. Beside me, an ice-cold glass of Coke will quietly weep its noble tears of condensation.)

Alexander Gorny · InKind · Original post ↗

Dining at a discount

The typical cost of ingredients in a restaurant meal is about 30 percent. Theoretically, it's profitable for an establishment to feed guests even at half price.

Besides, a restaurant is a typical small business. It's always just a step away from a cash crunch. Banks don't give them loans. Money now is much more important than money in the future.

The American startup InKind combined these two facts into a single product. The startup gives the restaurant an interest-free loan that doesn't need to be paid back. But there are no miracles—to pay it off, the restaurant has to feed the startup's guest-clients, naturally at a discount. A person eats for 100 dollars, and 50 is written off the debt. The ingredients, as we recall, cost 30—so in theory, the restaurant comes out ahead.

For the guest, InKind offers a "pay at restaurants through us, get cashback" mechanic. From that same hundred, the user gets 20 bucks back—a great offer, no reason to refuse. As a result, out of thin air, everyone in the deal ends up in the black: +30 for the startup, +20 for the customer, and +20 plus an interest-free loan for the restaurant. At least, that's how the pitch deck looks.

The obvious real-world problem is that the guest might not be new, but an old one. Previously, they paid full price, but now only half comes from them—for the establishment, the plus abruptly turns into a minus. The startup, of course, promises to prevent this situation with Smart Algorithms and Proper Recommendations, but there is no data. However, maybe as long as there are relatively few restaurants in its network, there is no problem. If you can only get a discount in one place in the entire neighborhood, people really will go there more often.

The project was launched in 2014 and has since developed without major public investments. But in a recent round, it raised a whopping 450 million dollars—most of it, apparently, simply as debt money to hand out as loans to restaurants.

Company website ↗

Bro, you're a real full-stack developer who can write anything independently. I think the way you use a model has nothing in common with the way I do.

For me, the strongest model isn't a top-tier piece of gear. It's the only way to complete a task of complexity x+n, where +n is precisely what only the strongest model can do out of the box today.

5.4 solved some problems of mine that were simply beyond 5.2's reasoning. And 5.2 was a leap into another universe compared with 5.1, especially 5.1 Codex Max.

I couldn't have made anything that works for me today on 5.1 Codex Max, even if I'd spent three lifetimes on it.

I've now rewritten my whole system with 5.4 xHigh, conducted by GPT 5.4 PRO. There were so many errors in the previous implementations that it would be easier to list what worked properly.

We're talking about a system of 474 files at that point. Some reached 32,000 lines.

5.4 refactored it all, and now I have 720 files. Everything really works differently now. It FLIES.

I used to see far more bugs, and everything was slow compared with how it works now.

Guys, as usual, I'm getting a fundamentally new experience through a game. Back in December, I made chess. That taught me about interfaces, dynamic elements, traffic, and plenty more that I then put to good use in my Salesbomb automation system, now at 230,000 lines including tests and failsafes.

Now I'm doing something completely new to me: working with a software distribution, file libraries, and a client being rewritten from scratch for my task, rather than the web over HTTP.

It's fucking mesmerizing to watch Codex juggle files, test servers, and builds. New versions of the client keep opening. Codex measures something, runs around a little, clicks buttons in the game, grunts in dissatisfaction, kills the process, rewrites something, recompiles, makes a new build, and tests again.

This goes on for hours. It works much longer with a live filesystem than with web services. It runs so many fucking tests that my computer has come alive. Development tools get installed, programs restart, and the reasoning monologue in the chat is completely different. It's genuinely digging around like a detective, openly pleased with progress, disappointed when it crashes again.

Such a fucking joy.

We dropped by a live interview with Bryan Cranston in the evening.

The guy performed All My Sons 118 times in four months (?????????). The last performance was on Saturday. His birthday was the same day.)

I've been saying for ages that Cranston is my favorite actor. Yesterday I understood why.)

I really hope the interview ends up on YouTube and you watch it too.

I've been to plenty of these live conversations with genuinely huge performers, McConaughey, Reynolds, and so on. But I'd never seen an entire audience stand at the end and applaud for several minutes, with tears in their eyes, shouting and whistling.

Cranston is a deeply philosophical guy. There's no question he answers without a parable, and at seventy he literally freestyles them. It felt like reading a fascinating book, every little corner of his reasoning so beautifully expressed.

And of course, he explained to me that I'm not some fucking talentless donkey wasting his life. I'm moving toward my dream in the only possible way. Bryan said he didn't know any great Hollywood star who'd earned their fame and money without working their ass off all the time.

He put it in a funny way: if you have no talent for what you're doing, you can skip the hard work, because it probably won't pay off. Better to find what you are talented at, or at least what you love doing.

But if you have talent, it'll only bear significant fruit if you work your ass off absolutely all the time. Because that's how the mathematics of society works.

“Some people hope or even believe they can achieve by healthy balance, but they will need to outwork me and my other friends from Hollywood who work all the time. They might bet on being genius but we are not stupid as well, and I don’t remember seeing genius interested in relaxing, genius is obsessed by applying the talent”

Reminder: the guy performed 118 fucking shows in four months :)))))))))))))

Here's what I'll say. Like most of you, apparently, I fell for 5.3 Codex xHigh because it genuinely was a leap forward and upward in fast coding. It probably was brilliant from every angle for the first one, two, three, four, or five days.

But a reasoning model is a reasoning model, and a coding model is a fucking coding model. Most likely, every model is more reasoning than coding for the first few days, then the weights start getting redistributed.

The stupid peasant on my left shoulder whispers, “But it says coding! You're coding! Use fucking coding for coding!”

The imaginary Einstein with dreadlocks smoking a crack pipe on my right shoulder, which is allowed because he's imaginary, says sensibly: “Do you think development is code? Fuck no, bro. It's lots of things, and you have to keep all of them in mind, see them, reason about them. Otherwise you get a sniper trained to shoot in a vacuum who wasn't told about wind. He shoots better than anyone but never hits when the wind blows.”

As usual, after wasting a couple of weeks, I snapped out of the spell of this bullshit marketing and put five fairly substantial tasks into 5.2. Not planning or a quick bit of reasoning, but “do the whole thing, sister.”

Slowly, at length, but so fucking well. What has been taking ten passes for a week and a half now takes one.

Quality matters so much more than speed when the difference isn't a year versus a month, or a month versus a week, but an hour versus five minutes.

Alexander Gorny · Sapiom · Original post ↗

Robot Access

Imagine a future where vibe coding has actually become a part of daily life. “Every Saturday at 7:00 PM, send me an SMS saying 'buy flowers for my wife,' and rotate the recommendation in a cycle: roses, gerberas, lilies, hyacinths, chrysanthemums, orchids. Occasionally mix up the order a bit so it looks unpredictable. If I'm on a business trip (check my calendar), skip that week.”

Claude goes whoosh-whoosh, writes all the code, but in reality, nothing works—you need to pay for hosting and for sending the SMS, and the AI doesn't have access to a credit card. Because if it did get access to a credit card, the owner wouldn't be thinking about flowers anymore; something would definitely go wrong. And you can't run it on a local computer because it gets turned off, and besides, real-world tasks might not be such toy examples.

The solution is offered by the American startup Sapiom—a unified billing system for all popular API services. A user links their credit card in one place, sets cautious limits, and grants access to their agents. Through Sapiom, they connect to web search, other AI models, or user authorization via SMS, and actively use them in their code. The code turns into a real product, and the world becomes a better place.

In a recent round, the startup raised 15 million dollars in investment, the first significant round in its history. Presumably, they will use this money to add more different APIs to the system. Right now, the flower task is still unsolvable: you can't send an arbitrary SMS, and Twilio or any equivalent is missing from the list of providers.

Company website ↗

Alexander Gorny · Plug · Original post ↗

The Boundaries of Niche Markets

People resell baby strollers to each other on Avito (Russia's dominant classifieds site) or OfferUp in the US. Drills are also resold on Avito or OfferUp. And flower pots too. None of these categories require a dedicated tool. But for used cars, besides the section on Avito, there are already specialized platforms like CarGurus or auto.ru (a major Russian automotive classifieds site)—the market size allows for a niche marketplace, and the specific nature of the product justifies its existence.

At the same time, while I'm not sure about Lamborghinis, Fords, Toyotas, and even Ladas (the ubiquitous Russian car brand) are resold in the same place—they are all the same, and there is no point in separate platforms for either businesses or consumers. But that's how it is now. In the future, Tesla and BYD will get their own independent marketplace—at least that's what the American startup Plug promises.

To the user, they explain the need for a new website mainly through tracking battery health. They claim, 'We know this is the most important thing in an EV, and we look at it closely, whereas none of our gas-powered competitors understand such nuances.' Honestly, this sounds highly doubtful. To investors, it's probably pitched differently—the market is growing, and we need to capture it while it's small. Right now, becoming the leader is relatively cheap, so let's become one, and in 5 to 10 years, it will pay off handsomely.

The mechanics of the project are just like the Russian service CarPrice (an online used-car auction platform). The service inspects the car, buyers bid on it in an auction, and the seller gets the money. It's fast and convenient, but cheap—nobody is looking for a car for themselves here, so the price naturally ends up below market value by the reseller's margin and Plug's commission. Though right now, the commission might even be negative—with low volumes, the startup might actually be paying extra to get things moving.

Over its 3 years of existence, the platform has processed $60 million in transactions. That's quite small, but still not zero; a few cars are sold every day. In its recent round, Plug raised $20 million in investment—the first big round in its history.

Company website ↗

February 2026

Fourteen months ago, the Christmas before last, we were visiting friends. Someone's sister's mother's husband, I forget whose sister or how we were all connected, was a British forensic specialist with fifty years of experience.

I told him about GPT. I think 4o had just come out then. Or maybe it was still 3o. He laughed and said it certainly didn't know anything about forensics.

I switched on voice mode and asked it to talk about the most advanced aspects of forensics, specifically in the UK and beyond.

He talked to it for five hours, the whole time we were drinking there. He gave me back my phone and said, unfortunately, this thing knows much more than he does.

Alexander Gorny · Span · Original post ↗

Saving on the Meter

As we all know, the sun and wind are free, and there is no cheaper energy. Take California, for example, where two-thirds of the electricity is already generated from renewable sources, and a kilowatt-hour there costs a mere 30 cents, compared to a whopping 18 cents on average across the US, or 12 cents in Louisiana, where there is almost no green energy.

For a household, this adds up to $100, $200, or $300 a month—painful even against the backdrop of California salaries, especially since not everyone living there is a Google engineer. But there is a way to ease the pain. Span is a smart electrical panel that manages your home's power grid and protects you from overpaying.

There is no magic here; an air conditioner won't consume less energy just because you have Span. But the panel skillfully juggles power sources to make every kilowatt-hour cheaper. It directs cheap nighttime electricity to charge your EV, stores free daytime power from your own solar panels in a home battery, and tries to avoid buying expensive evening grid power altogether—after all, that battery wasn't charged during the day for nothing. In real life, the scenarios are more complex, of course, but that is the general idea.

The price of the Span panel itself starts at $2,500, with installation costing extra. If it saves you, say, $100 a month, the payback period is about three years, not even factoring in future electricity rate hikes. For some, this might actually be a rational purchase.

The startup itself is doing quite well; it recently raised $163 million in funding.

Company website ↗

A couple of months ago I asked the community what to use for presentations. People suggested various things.

Now you can simply make presentations with 5.3 Codex xHigh in the Codex app. A document like this takes two prompts.

The first prompt read the brief and built a skeleton using a design from a popular library, which I'd also turned into a skill with one prompt last week, plus the PDF skill from Codex's default skill store.

The second prompt checked alignment and corrected all the text to what I needed. That's much easier to prompt once you can see what you're correcting.

Previously I'd have spent a week fucking around with a file like this, and it would have looked much clumsier. I'm completely hopeless at presentations.

What's the fucking point? Every presentation tool I've tried allows, at best, a teeth-grinding compromise. Or it's fine if you don't care about the presentation. Now you can make something ready to use for real work in Codex in an hour. Why use any other service?

Speaking of tourism, since it came up.

Once, Khun Raiwat invited me to dinner at his home. Khun is a sort of highest form of respect in Thailand.

Raiwat was Phuket's top politician. After the military and government reforms that happened while I was living in Thailand, he became one of the five most influential political operators in the whole country.

Here's one way to describe him.

I once flew back to Thailand from a business trip for the who-knows-how-many-th time. I'd filled more than one passport in two years there, so it was probably in the hundreds. Suddenly, the guy at border control decided to show off and told me I wasn't entering the country this time. I'd been coming far too often, and they didn't need tourists like that.

I spent quite a while persuading him to let me in. After a twelve-hour flight with a hangover, the thought of having to solve another problem was literally killing me.

Eventually, the Thai officer switched from “fuck off entirely” to “let's figure out where you live here.” Out of habit, I gave him the address: Soi Aree 1.

Aree is Raiwat's surname. Number 1 was his first luxury residence, which he later rented to Kostya Kalinov. It became the Aviasales office ;)

When the officer saw the address on the paper, he looked up and asked, very frightened: “You live at Khun Raiwat's?!?!!?!?!” I said, “Yeah, why?”

He ran out of his little glass cave, grabbed two other border officers by the hand, and they lined up and saluted me. The guy who'd been giving me shit started apologizing loudly and incoherently, assuring me it would never happen again =))) I had to remind him to stamp my passport. Otherwise I'd have had more fucking trouble on the way back, only this time I'd have had to call Raiwat's daughter too.

Anyway, there I am at dinner, surrounded by Raiwat's family, everything Thai-style luxurious. He talks to me about food for half an hour, then explains why I'm there.

“They tell me you know about various aspects of travel, including hotels. I have lots of hotels and condominiums, and I'd like to take advantage of the situation and rent them out much better with your help than the staff of my management company have managed.”

So I gave them a crash course in Booking and the then-emerging Airbnb, tossed out some ideas for extra affiliate revenue, and Raiwat was impressed enough to have a drink with me. His daughter said that was a great honor.

We were just shooting the shit after that, when I blurted out something about Thailand's tourism revenue and how important friendship with Russia as a country must be to them. Raiwat froze, narrowed his eyes, then burst out laughing.

“Fuck, Khun Anton. Of course you can't know about all this, so allow me to give you a little consultation on Thai politics in return.

“You have no idea how little we GIVE A SHIT about Russian tourists. The only people we care less about are all the other tourists. What share of our GDP do you think comes from tourism?”

I pulled something like 50% out of my ass, secretly thinking it was probably more like 85. Raiwat roared with laughter, spilling whiskey from his glass.

“Not even 8%, my young white friend. And that 8% is the most volatile, least reliable part of our GDP. You can't base policy on that.

“Next time you're traveling the world, look at what's written on car tires. Where are they most often made? Look at where the cars themselves are made, the ones driving around a huge part of the world economy. Eventually you'll learn about Thailand's medical sector, agriculture, all the different areas of light industry. And you'll never again think of insulting a Thai politician with the very idea that we might make our money from some embarrassing little trifle like tourism.”

I had a call today with my good old friend from back in Novosibirsk, Vanya Yagoda. Fuck, the man lucked out with that surname, which means “berry,” especially considering he's a full-time artist =)

When we were young, Vanya had two partners: Ivan Fans and Marina Yagoda. Marina and Vanya got married along the way :) They had a graffiti crew called TAKNADO!, very well known within the small world of Siberian hip-hop.

They did everything they could to brighten up gray Novosibirsk. When I remember childhood and youth there, through the fog of memory I see not just pieces of the city but their paintings in the background.

It wasn't just a few good drawings. Seriously, there was a time when people in Novosibirsk arranged to meet not at a street address, but at a particular wall painting or an object decorated by TAKNADO!

Even ten or fifteen years later, as far as I know, there are still objects considered neighborhood treasures in various cities. Like the tanks painted to look like condensed-milk cans in Yekaterinburg.

The most famous Novosibirsk rappers back then dedicated tracks like this to them =)

https://www.youtube.com/watch?v=tEYga649gMw&list=RDtEYga649gMw&start_radio=1

For the last four years, Vanya and Marina have lived in Spain, in the garden town of Estepona. Which, it turns out, hosts a huge international street-art festival every year =) Funny that they landed there by coincidence.)

I'd wanted to act on a fantasy I never get around to because of this fucking vibe coding: raise a pot of money for some profound, epic piece of art they could make for all of us to enjoy. Something about the pain and exhilaration of the present day.

But Vanya said he couldn't dive into anything like that just yet because... he sat down to vibe code in November and still can't tear himself away.)))))))

He made this:

https://personalia.art/en

For a token amount, you can generate an AI portrait of yourself and have it delivered by mail in stylish, beautiful packaging.)

So far, the entire flow for men is finished, and Vanya is now working on the women's branch =)

A nice gift for someone who doesn't drink, when you can't get away with a bottle of whiskey.) I usually give Soho Home glasses in that situation. But you can't give glasses twice in a row. You need more ideas, and this is one.) Click around if your finger has any energy left by evening. If you have feedback, I'll gratefully pass it to Vanya.)

Conceptually, it's cool to build something where AI meets the real world. Vanya is getting some initial traction and plans to find investors for the next stage. I hope it works out and makes it easier to get AI art of any kind onto your walls and shelves.

Here are Vanya's and Marina's Instagrams, by the way. Beautiful stuff =)

https://www.instagram.com/dondonberry/ https://www.instagram.com/marinayagoda/

And here are lots of videos about different pieces from the last fifteen years. The very first, oldest video on the channel is about the work that taught a lot of Novosibirsk kids what funk was. Because sometimes it's better to see something once than...

https://www.youtube.com/@Taknado

It's all possible, but it isn't easy, to put it mildly. Building Google is possible too. Sergey Brin built it, didn't he?

All I can say is that the amount of headache involved is such that I haven't seen ANY successful examples outside Salesbomb. You'd think I might have encountered one in twelve years.

A couple of times, I found people who came recommended and looked impressive. But on closer inspection, they turned out to be scammers.

In my world, a scam isn't necessarily outright stealing money. It's taking money to perform a set of motions where a result could happen only through some extraordinary, epoch-making miracle.

In our business, you have to bring thousands of different parameters together. Then survive the level of market resistance.

Nobody likes lead generation. Everyone wants to kick it in the fucking face. With everything that follows.

Most businesses can't even immediately tell you why they're surviving. Mostly, it's founders running around in a panic and somehow making things happen.

There is an alternative to lead generation. Most companies don't have it, yet they exist, pay salaries, and grow. Most often, the missing ingredient is market fit plus a business model plus the right fundraising.

If all three are good, it's like sticking a pole into the ground in Texas and having oil gush out.

The others usually hustle through friends and family, land one big client—in Russia, usually some big bank like VTB24—and pray. Sometimes prayer works for decades.

Criticizing anyone is pointless. If a guy currently has money to travel, eat at restaurants, and get into a new car, good for him.

Alexander Gorny · Balance · Original post ↗

Renting Instead of a Mortgage

It's a classic financial distress storyline: a person takes their belongings from home to a pawnshop, then they run out of things, and after that, some real nightmare begins. The American startup Balance offers to add another intermediate stop to this route—you can carry the house itself out of the house!

The startup pays off the client's mortgage debt to the bank, hands them tens of thousands of dollars in cash if they want, and converts the total amount into an equity stake in their home. This stake now belongs to Balance. From that day on, instead of making mortgage payments, the user pays rent for the third or half of the house they just sold. Typically, this is a significantly smaller amount, making life easier.

Once the person climbs out of their financial hole, the process goes in reverse: the share in the house is bought back using a new mortgage from a regular bank. The buyout price is the original value multiplied by the change in the home price index over the elapsed time. Aside from the losses on commissions, it's a pretty fair deal.

In a worst-case scenario, nothing changes. Then, after 7 years, the contract with Balance expires, and the client faces a choice: buy it back or sell the rest of it completely. Most likely, the majority end up selling.

For the startup, this entire mechanic is a classic real estate investment. The purchase is made at a discount, as the client is in no position to haggle. The rental stream is almost guaranteed—where is a person going to go from their own home? On top of that, Balance receives a transaction fee instead of paying one.

The startup was launched in 2021 and managed to get acquired as early as 2023. In 2024, its acquirer, EasyKnock, shut down, and this January, Balance announced a new $30 million round and a business revival.

Company website ↗

Before bed, let me share “how not to do it.”

I went a little crazy about development quality. What if I made the machine take one tiny task at a time toward a goal made up of hundreds of subgoals, where 100% means a system that fully works and has been tested from every angle?

I implemented that setup and made a looping prompt that was supposed to build 100% of the system.

But after 300 (!!!!!!!!!!!!!!!!!!!!!) commits over forty hours, I started to feel we weren't going anywhere anymore. I launched an investigation. The system was so obsessed with the quality of each iteration, exactly as we'd asked, that it had deteriorated into endlessly rewriting the same piece. Over 180 commits, AI added just 300 lines of code, and 100 of those lines were rewritten more than a hundred times.

I tried to fix it, without much success.

It seems more effective to build a shitty system quickly, then refactor it into something reasonably good, and finally polish the specific pieces that need to be perfect.

Otherwise the model simply gets stuck in a death loop of its own strange, impossible goals.

Alexander Gorny · OpenEvidence · Original post ↗

Medical Chat

Doctors are a lucrative target audience. You can show them highly expensive ads, and they control massive budgets that aren't theirs. Whatever medicine the doctor prescribes is exactly what the patient or their insurance company will buy. It's a perfect setup.

AI is the trendiest topic of the decade—no explanation needed there.

The American startup OpenEvidence combined these two facts and created an AI for doctors. Its neural network is tailored for medical questions and has access to a library of the latest research. It is supposedly better than general-purpose LLMs at answering questions about drug compatibility or unusual symptoms.

It looks like a regular chat, and access is free, but only if you scan your medical license—outsiders are kept out. However, a test version is left exposed to the public, and it did answer my question—and, interestingly, in a completely different way than ChatGPT. ChatGPT comments, 'most likely, there is no problem,' while this one immediately suggests rather aggressive treatment methods. Maybe it was just a coincidence (based on a single query, after all), or maybe the difference is that a patient is more likely to like the phrase 'you are healthy' than a doctor is.

OpenEvidence claims that 40% of American primary care physicians are already using the system—a massive number for a relatively young project. Pharma has also noticed it; the startup's current revenue is around $10 million a month, and it recently raised $250 million from investors at a $12.5 billion valuation. All in all, things are looking great for the project, whichever way you look at it.

Company website ↗

A friend upset me yesterday. He says I'm coding all wrong. What I do in fifty hours can be done in three if I use different agents working in parallel.

I really don't parallelize processes, because I tried twice. Both times, the coding was faster, but cleaning up the problems took me ten times longer. Even if agents share some source-of-truth context in an awkward little file, in a burst of creativity they build something the other agents don't know in detail. Those differences in detail are exactly where everything goes to shit.

But I last tried it last year. Maybe something has changed since then? Please talk about this a little!

Alexander Gorny · Full Day Handyman · Original post ↗

A Handyman for the Day

One of the largest consumer markets where marketplaces or Uber-like services haven't won yet is home repair in the US. Fixing an outlet or a dripping faucet is still a nightmare, a pain, and a completely unreasonable expense. Naturally, over the last 20 years, a wide variety of services have been launched there, but nobody has found 'that exact model' yet.

Another attempt is Full Day Handyman, a startup whose name completely describes the concept. A handy young guy comes to your house for the whole day and solves all your accumulated problems—hanging things, fixing things, and so on. This costs a thousand dollars per visit. This is more than enough for the contractor to make a decent living, but significantly less than if you were to estimate the market value of each repair individually.

As a result, customers save money, workers are happy, and estimators along with a bunch of other redundant middlemen are left out of the game—just like in the best examples of the Uber economy. However, this is still largely theoretical. The idea for the project only came up this past December, and it is taking its very first steps. The key mechanics and customer acquisition costs have been tested on only a few dozen orders. For those, the results looked very good on both sides of the marketplace.

Right now, the project's founder is looking for a co-founder partner—an online marketing genius. There is a chance to build a unicorn from scratch. Besides the obvious experience requirements, it is important to live somewhere in the Americas—the time zone matters. The founder himself is Roman Levitsky; in Russia, he ran the Ruport advertising agency—the very agency that once managed to bring Elon Musk to a business forum in Russia's southern Krasnodar region. If you are interested in the project, message him directly on Telegram at @RLevitsky.

Company website ↗

The devil really is in the tiny details of “how to ask properly.”

You can say, “Rely on the file where you recorded the previous action.” Or you can say, “Rely on the file where you recorded the previous action, but ONLY if you've rechecked all the code written then and what actually appears in the logs and database after the function runs.”

That's the difference between bias with two minutes of reasoning and forcibly discarding that bias with twenty minutes of reasoning.

That's why you have to THINK THROUGH these looping prompts for building things. When I make a bad one, I can tell immediately from how quickly the model spits out results. Give it twenty prompts, and it's finished in thirty minutes.

With a proper prompt, the model physically can't skip the rechecks. Those take most of the time.

Now I spend about thirty minutes before bed writing a looping prompt in the direction I need. I wake up, and it's still working.

Sleep has become more enjoyable. Important, difficult work gets done while I sleep. That's fucking wild.

During the day too. It frees me up for other work and tasks. When the average prompt finishes every ten or twenty minutes, you switch back so often that you feel as if you've spent most of the day coding instead of doing other things. It probably isn't just a feeling.

But if you've thought through where we're digging, written the looping prompt, sent it off, and know Codex is busy and unavailable for the next eight hours, you can fucking start living instead of wasting away staring at the terminal.

Today I'll try setting Codex on building a complete Instantly clone without the customer-facing functionality, just for our agency's use. I wonder whether giving it the task that way will produce a real, full-fledged Instantly in about fifteen hours.

If it does, you could build all the software you use or want to use in a week.

Another thing I like about chess: it shows me where my mind is on its current wavelength.

In almost two and a half years of playing, I've noticed something interesting.

I've always paid for premium on Chess.com. At the end of a game, it shows me the accuracy of my moves and breaks it down. That number is far from exact, but it doesn't need to be exact for what I'm talking about. Seeing the relative difference is enough.

Everything in the world consists of harmonics, of oscillations. Even string theory is called that for a reason: everything in the world is a string with some amplitude.

Our personality, intelligence, brain, and so on consist of countless variables. But experience tells us that sometimes coming up with things is easier, and sometimes it's harder. Sometimes the words flow eloquently. Other times it's “God, let me manage to formulate a thought at all, and that'll be enough.”

Chess and its accuracy score made that process visible to me.

If I play every day, even just five or ten one-minute bullet games, I can clearly see my brain's current tuning for whatever kind of activity matters in chess. I have no idea what kind of activity that is. Logic, calculating ability, mathematical intuition, or all of it together.

There are days when even in the most chaotic games I get at least 78–82% accuracy, often over 90% in short games.

Then there are days when I think I played pretty well and the games seem interesting, but the accuracy reads 30–55%. I'm simply incapable of playing any better that day.

Trying harder doesn't change that. Effort helps me reach the top of today's wave. But if today's wave sucks, then tomorrow, or the next day, or a few days later, I'll play 20–30% better without any effort than I can today trying my absolute hardest.

Today, incidentally, I'm beautifully dialed in: five games in a row, all between 82 and 89%.

It doesn't directly depend on stress, sleep, or food. From everything I've observed so far, in my case it seems to depend most on what's happening outside the window: natural phenomena, changes in the weather, and so on. I've had an awful day, full of fucking enormous stress, and I feel like a stupid asshole. But chess says otherwise.

The day after tomorrow I might have a magnificent day, feel full of energy, feel capable of anything, and chess might say the opposite.

And that doesn't mean I'm a stupid dickhead that day. On days when I'm terrible at chess, I often do something else exceptionally well, by my standards. None of this really means anything. I'm like a singer from Yakutia, up in Siberia: whatever I see is what I sing about in this long, lustful song of mine.

That's a great question, and it turns out the answer is very simple.

I knew all this automatically because it's my life. But when I started building a system to manage our ecosystem, the problem became clear.

People think: fuck, outreach is outreach. You send emails. That's it.

Contacts. Emails. Money.

I realized how complicated and fucking crazy this little world is when I sat down to automate the agency. Actually, not when I sat down. Two weeks later, when it turned out I'd barely even started.

I don't have over 150,000 lines of code because I love fucking code. I'd rather code a game or something else. I have plenty of desires much sweeter than the ones I'm working on now.

But mailbox orchestration alone takes THREE separate pages in our control panel, fucking packed with features, triggers, and monitoring.

That's the main reason people get nowhere with outreach. They don't see the iceberg.

Yesterday I asked who knew anything about UI/UX. A kind person shared these skills.

I asked Codex to scan both links for security risks. If everything was okay, it should understand both skills, install the better one, test it, and, if successful, apply it to my project. In one prompt, everything was actually checked, selected, installed, and the design completely rewritten. Codex said the skill at the first link won unconditionally.

I noticed a benefit of using skills: the task finished several times faster and used several times fewer tokens. All the operations took 48% of the very first context window, without compaction.

If you don't have a UI/UX skill installed, I recommend it. These are the skills a friend recommended in the private chat:

https://ui-ux-pro-max-skill.nextlevelbuilder.io/

https://github.com/vercel-labs/agent-skills

I liked this pack too; I'll try it this evening. You can point it at the shadcn library. Most SaaS products seem to be built on it these days.

https://ui.shadcn.com/

Yesterday the fucking gorgeous Olya, an actual supermodel, suddenly decided to host a little gathering at her apartment. Two absurdly fucking smart guys from the British academy of sciences gave talks: a physicist on string theory and a mathematician on quantum processors.

The most popular answer to our questions went something like: “Fuck, I could explain it, of course, but you wouldn't understand... Okay, look. If we take Fourier series and fungicize compound two-factor molecular bases into interpolar forzaginal transgricification of material atrophy, then in the gravitational vortex of intercongensualization of protomorphic crystals of weakened-proton dualization, a curious photonic distortion of asdfghjkl properties appears...”

But the cake was so sweet that the sugar boost seems to have made me understand everything anyway. It's amazing they managed to seal that much fucking sugar inside such a tiny piece. I haven't pooped since, and I don't think I ever will again.

For some reason the British can't make cakes any other way. They all graduate from Oxford and get PhDs, so I think this is their kind of laboratory showing off. “Fuck, I packed three to the nineteenth power of sugar into one cake molecule.”

“That's nothing, buddy. I've been packing three to the twenty-first into a cake atom for ages. Two hundred more people to kill and I unlock the twenty-second power. Loser.”

The Codex app has one significant drawback: it dies.

Has anyone encountered this and investigated what the fuck can be done?

Once a day, UNTIL I RESTART THE COMPUTER, the interface stops showing any updates to what's happening. Even closing and reopening the app doesn't help. Well, it refreshes the picture of the tasks, but if you send another prompt, it freezes again.

It's a really stupid problem. The app's frontend freezes, while the backend finishes its task normally.

I've never even encountered something like this. I don't want to go back to the terminal. Help me fix the Codex app!

It feels as though something's cache fills up and can't clear itself. But I haven't found any settings that even remotely control it.

Yeah, I didn't bother writing about it, but Codex eventually dug up the problem: a new parent server gets created from time to time, for no obvious reason, and the old one doesn't switch the interface over to the new one. So Codex made me a script that does it every 60 seconds. It also clears the log file, which had reached two gigs. No idea what the fuck it was for, since no LLM is going to read two gigs of logs anyway. Even the idea that I ruined my experience because Codex wouldn't catch things in the background is nonsense: we still let the log grow to 500 megs.

Now Codex flies. It used to respond slowly to prompts being sent; now it responds in a millisecond, like the terminal.

The app is probably still rough around the edges. Eventually I'll be able to kill this worker, but for now it's fixing OpenAI's developers' mistakes itself.

Today I tested what I threw together on 5.3 before bed yesterday. What can I say? Another breakthrough. It's become a hundred times more powerful.

The evidence isn't just that five difficult tasks worked perfectly on the first attempt. It now refactors adjacent areas on the fly. What 5.2 considered the best solution, 5.3 considers a workaround and poorly optimized crap code.

With one request, it rewrote my orchestration of API requests to external systems without breaking anything. Load dropped twentyfold, speed improved noticeably, and the number of errors in request results fell to zero.

5.2 had optimized the same area before, but did what it did. The ceiling of 5.2's reasoning in code is below the floor of 5.3's reasoning.

That's how it is.

Guys, last night, just before bed, exhausted and wrecked, I inexplicably pressed a couple more buttons on my phone.

The browser slowly read my query, syllable by syllable: “vibecode via codex on mobile phone.” Google's first result was the landing page for the brand-new Codex app. Stupid little bastard: there isn't a mobile Codex app.

But I clicked anyway and decided to read the page. Earlier that day I'd only read the Download for macOS button. I found a section called “Automations.”

At first I didn't understand. Then I fucking UNDERSTOOD.

This is science fiction. You can set up prompts that launch themselves on a schedule?

That means automatically debugging logs, minor refactoring, a couple of one-off tasks that didn't get done during the day, the tasks you least like doing, and so on.

WHILE YOU SLEEP.

I'll see what else this automation can do. I think there was more, but I forced myself to sleep and stop reading. Skills are great too, by the way.

You can now build a trend-based TikTok video generator entirely in Codex without even breaking a sweat.

January 2026

Alexander Gorny · Erebor · Original post ↗

On Big Rounds

We've somehow gotten used to the idea that a couple of brilliant research engineers can leave to launch their own AI startup and immediately land a valuation of a couple billion dollars and a couple hundred million in funding. If Ilya Sutskever helped build OpenAI, then his new project must also be worth a fortune—or so the logic goes. It's hard for me to accept this logic, but easy to understand and remember. That's just how the world works; they've even coined a special term for these kinds of rounds. In any case, we can only congratulate Ilya.

Now, this has started happening with non-AI startups as well. The American startup Erebor is a typical bank for tech entrepreneurs, a direct analog to Mercury or Brex. There is, however, one major difference from its competitors: Erebor doesn't actually exist yet. It has no clients, no website, no app. Nothing but a name and a license.

And just the other day, this startup raised $350 million at a $4 billion valuation. Four billion dollars for nothing. Brex is already a massive company, a market leader, and was recently sold for $5 billion. A year ago, Mercury raised funding at a $3.5 billion valuation. Erebor, with its "nothing," is valued roughly the same as the companies it would need about five years to catch up to, assuming everything goes perfectly.

Of course, the startup's founders are highly respected, to say the least: one built Oculus and Anduril, the other Palantir (and yes, apparently they are both Tolkien fans). And they probably wouldn't even get off the couch for any less money. But I still don't understand how VC funds see the math of this investment. In theory, yet another AI startup could eventually be worth a trillion dollars, which explains the bet on it. But another bank for startups?.. How?..

Company website ↗

Every week of vibe coding makes it clearer why things never worked out for me with actual flesh-and-blood development teams.

It's about a way of thinking and whether it fits within physical limitations. Having to write code manually with a small, budget-limited number of hands that get fucking tired quickly is a serious physical limitation, to put it mildly.

My way of thinking produces a fucking mountain of tiny features. They're decorative or supporting details that seem unnecessary, or whose purpose isn't obvious at first. But together, they make up the user experience of the thing I'm building.

About nine days ago, I took a seriously wrong turn. Yesterday, I concluded that rewriting everything from scratch would be easier than having Codex refactor a hundred thousand lines of code. Honestly, I'm not even sure 5.2 could refactor it.

When I realized I'd have to describe in words everything useful I'd built in there, I got scared. I sat down to do it anyway. Two hours later, I hadn't described even half of it and had completely confused myself. Each of dozens of features has eighty subfeatures that make a lot of sense together, although individually they make none.

Luckily, GPT 5.2 PRO can produce a structured map of features, with a short description of each and an explanation of its logic and relationships with the others. Then you can simply delete everything you don't need and be left with exactly what you do.

But I had to fuck around with that for a full twenty-four hours without sleep.

Which is interesting too. Before, nobody expected anything from a day. Or even a week. Now a wasted day feels like a genuine Shakespearean tragedy.

I think the main thing that will trip up most people getting into vibe coding is persistence.

You don't need persistence just to crank out loads of features. The kick you get from each one gives you enough momentum to build the next two.

You need it when a project keeps getting more complex and you hit a problem that requires banging your head against a wall before you can solve it.

I've just overcome a problem that CODEX, GPT 5.2 PRO, and I spent… exactly thirty-five hours fighting. Over sixty commits, a billion different deep investigations of the whole codebase, external documentation, and so on.

The cause turned out to be stupid. But it had hidden itself so cleverly that neither GPT 5.2 PRO nor CODEX could see it. I literally found it by trying everything one by one.

I've spent years looking at the texts people write for work and understanding why the quality is so low: people can't make themselves work on something that doesn't give them enormous pleasure for longer than a certain amount of time. So I also understand that not many people are willing to spend thirty-five hours on the exhausting business of hunting for shit in code.

That's exactly why I'm sure an autonomous agent that codes from beginning to end without human intervention, and tests until it works, will appear this year. Most likely, it's very close. Without that piece, most people will get disillusioned with vibe coding after failing to make anything substantial and truly complex even for themselves, let alone sell their solutions for money.

I wonder whether that thing will appear in March or April.

This is the first time I've encountered something like it.

Previously, that was precisely the problem: if collaboration between Codex or Claude and a person was needed, the process became almost impossible. A coding agent assumes nonstop motion. If you needed to complete an OAuth flow along the way, for example, you were screwed. You had to stop it, complete OAuth, then ask again.

What Codex has just done is fundamentally new. It actually worked through it with me step by step WITHOUT BEING ASKED. It even accounted for its own limitations and used them to suggest the best sequence of actions.

I'm curious when it learned this. It might literally have been in the latest updates.

It's the best real-life illustration of what several great contemporary thinkers have said: when AGI arrives, there won't be fireworks. We may not even notice.

Alexander Gorny · Bilt · Original post ↗

The Best Loyalty Program

1.5% cashback on socks on Mondays, 2% on cafes on Tuesdays, 5% on taxis on Wednesdays! Banking apps look like decorated Christmas trees, bonuses fly in from left and right, and you just have to grab them in time. But the massive gaping hole in all loyalty programs is rent. For many people, this is their single largest expense, and they would love to get cashback on it first and foremost—but, alas, there is nothing.

The American startup Bilt filled this frustrating gap. It came up with cashback for rent. A customer gets a Bilt card, pays the startup with it, and the startup takes the money and sends it to the landlord via bank transfer. Nothing changes on the landlord's end, while the tenant gets bonuses in their app—around one percent of the payment, amounting to tens or hundreds of dollars a month for doing nothing extra.

Bilt's monetization came from three sides. Landlords paid a small fee to connect their buildings. The startup claimed that for the sake of cashback, tenants kept to their payment schedules better and moved out less often—no one wants to lose their streaks and boosted rates over a single day of delay. In addition, the bonuses could be converted into programs of other networks. Those networks paid for acquiring new customers or, at the very least, allowed Bilt to spend less on payouts.

And also, a crucial role in the business model was played by the partner bank, Wells Fargo, which issued all the cards. Bilt somehow managed to strike such a sweet deal that its partner was allegedly sitting on losses of up to 10 million dollars a month—unpleasant even for a major American bank. At the beginning of the year, the bank rebelled, tore up the agreement, and the cashback ended. Right now, Bilt is urgently transitioning customers to different terms with another bank, but that will be a whole new fairy tale. Users generally perceive this relaunch as the death of the service they were used to.

Company website ↗

Today, in the pub's smoking area, Andrew and I remembered how fucking great Avril Lavigne was on MTV.

When I got home, I decided to stick a finger in the cassette of passing time and rewind it. I put her tracks on Spotify.

Memory is a winter-coat pocket you haven't reached into for decades, maybe, and then you find a 100-ruble note in it.

I saw that very sheet of notebook paper. I'm not kidding: I hadn't thought about it in twenty years. Had I been embarrassed by it? It appeared so clearly before my eyes that I felt I could reach out and touch the paper again, warm from my nervous palms. On it, I'd painstakingly written my message to Avril.

My scepter of love hadn't yet been tempered in the furnace of female enthusiasm, so I had no idea what to do with that feeling except try to describe it in letters. I thought if my handwriting was beautiful, she'd understand everything.

I ran around the apartment looking for “the right” pen. Then I went to Anya in the next building entrance. Her mother taught Russian language and literature, and a child's logic is simple: if someone teaches words, their home might contain tools that can make words even more real.

I don't remember a single thing from the letter. But I remember being nervous at the post office on the outskirts of Novosibirsk, pleading: “Please, just don't get it wrong. Put enough stamps on it so it definitely gets there, to America...”

There was no reply, of course.

But the letter arrived anyway. Not to her. To me, through this memory.

It arrived across the years, across thousands of “well, that happened,” across the repairs to the soul we call experience. It arrived and knocked: “Bro, do you remember who you were when you didn't know how to love yet, and loved anyway?”

The irony is that I've been doing the same thing ever since.

Writing into the digital beyond with passionate intentions. Writing to people who probably won't reply. Though sometimes they do, and that's the greatest joy =)

Another problem my colleagues and I regularly discuss is the number of parameters being considered.

We tell a client something. They don't understand, and one of us starts getting annoyed. But consider how many parameters we're touching on in that sentence and how many they are, and getting angry immediately becomes impossible. The difference can be thousands or even millions of times, even though we've said exactly the same words.

I've been in fucking sales practically since diapers. Even if I were a complete idiot, I'd probably have picked something up in thirty-eight years. And I'm not quite a complete idiot yet. Still learning.

Here's a simple example if that wasn't clear. Take one word:

PROFITABILITY.

Ask a schoolkid, an ordinary clerk—even one who works at a bank—and someone like banker Oleg Tinkov to write everything they can think of about that word.

Same word. You wouldn't even see the difference between the schoolkid and the bank clerk next to the epic Tinkov would spend three years writing.

That's why mentoring is such complete crap without years of putting it into practice together.

Honestly, I don't understand the question.

A company isn't static. It's always changing. What causes the changes doesn't matter. If a company can't adapt naturally to change without falling apart, it's a company that exists only on paper, or one that doesn't survive by selling to customers: a laundering operation, a hobby funded with spare cash, a sham built with stolen money, and so on.

If a company has existed for ten years and operates on its own money, repeatedly earning it through sales, you won't destroy it with innovation unless the innovation costs more than it can afford. Or unless it destroys sales channels or hurts the people doing the selling.

I do the selling, and counting money is an inborn function of mine. I don't engage in self-harm either. The team just needs to do its part of the work.

If I'm doing something wrong, they resist. If their position is proven, I back the fuck off. It regularly happens that I imagine something and they explain that it's nothing like that. But just as often I bring a great idea, we implement it, and everything's fine.

Maybe you meant something else: how people see my outrageous innovations that replace them. I don't give a fuck. If someone gets replaced, that's their defeat, not mine.

I'm a closer and I work with closers. Closers don't lose if the task is achievable. If someone isn't prepared to invent something so I can earn money and share it with them, and I've acquired the ability to stop gifting it to them, that means I won.

When you clean out a poker table, do you hesitate about taking the money? I don't. It's a game with specific rules. If someone doesn't want to understand the rules or is tired of playing, that isn't the problem of those willing to continue.

And I don't build “family companies,” where everyone is one big fucking happy family. I'm bad at hypocrisy. We aren't a family. Everyone gets something from working together, and when they've had enough, they leave without looking back.

That's simply what actually happens. I don't need all those sexy social manipulations. I have plenty of emotions inside my own head. More than enough.

But this is an agency business, like in Suits. If I were building a product company, I'd create a cult. I'd have to learn hypocrisy, but that's no harder than vibe coding.

Well, I'm no longer a LinkedIn Premium user.

“Subscription was canceled successfully. We're sorry to see you go.”

The audience is in tears.

I've spent $20,196 on LinkedIn over seventeen years. Not a single month without the $99 Sales Navigator subscription. Although Sales Navigator didn't appear immediately. For the first couple of years, it was a fairly convenient LinkedIn Premium. Only later did they completely lose their fucking minds and give the world four different interfaces, every one of them impossible to use.

Goodbye, LinkedIn. See you in fifteen minutes, when I buy a subscription with the new card you won't let me add without canceling the old subscription, you fucking idiot.

Unfortunately, LinkedIn is absolutely irreplaceable. Every pixel tells me something about a person that stops being so revealing when presented anywhere else. Even if I could get all that information for free, I'd still keep paying.

It's a little sad that, as usual, I use 0.01% of the functionality in a $99 subscription. LinkedIn probably doesn't even consider that part something it monetizes. But that's life.

I never even open Sales Navigator itself. Yet that's the bastard I have to pay for.

And I swear, the paid badge is DEFINITELY one of the ten most important things an LLM looks at when scrutinizing a job candidate's suitability, regardless of the position. I don't need that, but I'd suggest others think about it and keep the badge on their page while job hunting.

I also think we should be ready for the possibility that $200 plans are an investment program to embed AI development in enough people's digestive systems, then make money from them once they can't live without AI.

A reminder: OpenAI said in 2024 that agents would cost roughly $2,000–2,500 a month.

At any moment, a model could arrive wrapped in the right marketing: it's fucking brilliant, but because of the new infrastructure, your $200 now burns through the tokens in two requests. Something like that.

Remember this tweet when using models like royalty costs thousands of dollars rather than $200 as it does now. It'll definitely happen. The only question is when. In theory, it should be soon.

That's not quite how it works. Capital is earned in the future through investments made in the past.

For example, in 2023, there was serious uncertainty about what kinds of startups would deliver big returns five years down the line. What happened? People simply stopped investing.

Nobody gives a fuck whether you're junior, senior, mid-level, or whoever else. Try being senior without any money.

Money is the foundation of all the clever words people use in the market.

Right now, people are guessing who AI is hurting. Just the really dim ones, or the somewhat smart ones too?

But AI hits everyone simply because projects, teams, and business areas that would previously have received funding and customers won't get them now. You can be a billion times senior.

It isn't just investment. It's customers too. Previously, corporation X would at least have considered buying startup Y's product. Today, there's a blanket ban on all that. The entire startup crowd can fuck off, as they say.

I've noticed an interesting pattern that your brain and hands apparently arrive at after a certain amount of mileage in vibe coding.

I've started coding blind a lot. My brain understands how to break tasks down properly and has learned to remember what was done, and in what order, over the last 10–20 commits. It automatically plots a direction for what still needs doing. My imagination has learned to picture the resulting interface without looking at the actual interface. I do anywhere from three to 15 commit-and-deploy cycles without opening the system I'm building at all.

I wouldn't do this if it produced broken results. The skill is precisely that my intuition now picks up the risks of each specific request well enough, while tests are added for every new feature, sometimes several different tests at once. So if a new feature breaks an old one, I learn about it from the pre-deployment report in Codex, rather than from clicking around the product.

I think this is a new kind of skill that LLMs are cultivating in people. I doubt this was possible when people wrote everything by hand.

It really helps combine coding with other work.

A month ago I sat meditating on what Codex was doing and how, spending my time on it. Now I simply do my usual non-coding work and quickly add tasks to Codex as they finish baking in my head. Put it this way: I used to spend seven to ten hours of my own time on ten hours of vibe coding that produced something for me.

Now I spend half an hour at most on ten hours of vibe coding that produces something for me.

Test cases aren't needed because the LLM understands all the code and what it's for better than I do. It's enough to have a document in the repo that the LLM has to read every time, covering the basic principles: which MCPs to use and why, what to do if one of the OAuth authorizations expires, how tests should work in general, and so on.

If you simply ask 5.2 Codex xHigh to dig through the code and work out where tests are missing, you'll be surprised how many it knocks out in eight minutes. A hundred, easily.

Yesterday my buddy Andrew and I got particularly spectacularly hammered. He holds the valiant position of CTO at a British bank.

First we had some strong beer with a steak. And another beer while the steak was cooking, obviously.

Then Andrew produced a bottle of lively Primitivo, and that went down beautifully.

But some kind of negative alchemy happened between the two liquids inside us. I remember my tongue starting to stumble as the bottle of red ran out. Andrew still seemed like an energetic young champ, but he started dropping things.)

Naturally, Andrew and I decided we were just TIRED after such a difficult working day and needed something to perk us up. I immediately remembered my French friends telling me they always had a little whiskey after a good steak and wine.

One little sip, then two. My fucking tongue still wouldn't stop getting stuck in my teeth, but we didn't give up until we'd put away a good two-thirds of a bottle of whiskey.)))))

To say today was difficult would be to say nothing.

But I remembered Andrew saying to me at some point: “Bro, listen, I've had a look at what you're vibe coding. It's pretty strange that you're actually getting any of it to work. Log into Git and let me see what it looks like?”

I logged in and handed it over.

Andrew clicked from project to project, looked at the tests, at how the core features were structured. Then he asked me very seriously: “Fuck, you really don't have a programmer who's made any changes to this?”

“No,” I said. “Codex just banged it all out itself. I don't even fucking know how to read it. What do you think of the code?”

He scratched his head. “Well, fuck, brother. We basically do it the same way this thing wrote it for you. The only difference I can see is that some of your files have eight thousand lines of code. Real programmers don't do that because it's very hard for a person to read. But a machine doesn't give a shit, so it's absolutely fine.

“I know you, Anton. You're a fucking obsessive, so this probably won't work for everyone just yet. But I also know your near-zero understanding of development. If obsessiveness alone is enough to make things like this practically for free, then in six months you won't even need the obsessiveness.”

Andrew drew no further conclusions. He thoughtfully clicked the remote, put on some early MTV videos, and by then properly marinated in whiskey, we happily dissolved into them.

I've had some new thoughts. I'll need to update it.

I've been thinking of doing streams where I update certain things and explain what's changed: market conditions, for example, or new software that's appeared.

For monitoring remote workers' activity, vibe coding is an incredibly powerful tool. You can describe code with tests. Or you can describe people.

If someone has a set of applications they absolutely have to use to do their job, otherwise the job can't be done, and that's anyone who works at a computer, you can vibe-code a monitoring system much better than sticking a camera in their face. Completely invisible, too.

For example, I know with 100% certainty who uses which parts of my automations in the company, who doesn't, how often, exactly what they use, and in what sequence.

You can listen to what people say and look at what they actually do, then sit drawing conclusions like fucking Freud. You can even puff on a pipe packed with imported tobacco.

If you don't do that, you'll still be sucking on something, but it won't be a pipe.

Today we made a triumphant 900-kilometer dash back the other way, offended by the lack of snow. So the snow decided to kick the living shit out of us on the highway.) We crawled through slush for 300 kilometers on summer tires.

Good thing I'm from Siberia, where we all play with the handbrake as kids, because I had to catch the car sliding more than once today. One time I don't even understand how we didn't fucking crash. That was probably the Christmas miracle.

At least I now know that four-wheel drive, all the incredible electronic wizardry, and brand-new Michelin summer tires with their tread still completely intact are all bullshit in the face of snow on the highway. Only Brits, who won't go out in the snow without chains, believe in that stuff. They'd been telling me it was incredibly fucking reliable.

Well, at least I got some fucking great drifting in today.) And the car's in one piece.) And how fucking great is it when there's a McDonald's next to the hotel! I've never been this lucky before.)

Alexander Gorny · Curative · Original post ↗

A Discount for a Checkup

A person with health insurance is tempted to practically live at the clinic—after all, why not exploit a free service? Insurance companies hate this, so in response, they invented deductibles and copays, where the patient pays part of the cost out of pocket. Not enough to go bankrupt over a broken arm, but enough to keep you from going to the doctor for no good reason. In the US, this is standard practice.

The result is that people do indeed save their money and avoid using medical services unnecessarily. But, alas, sometimes they save too much and don't use them when they actually should. Stomach hurts? Oh well, it'll probably pass. No sense in paying for an exam and tests. And then, when things get critical, it would be one thing if they just died—but no, they can be cured, just at a much higher cost. The insurance company is in the red again.

The American startup Curative comes full circle and brings back real, comprehensive health insurance with zero out-of-pocket costs for the patient. There are two conditions: only visit clinics from a pre-approved list with whom the startup has negotiated good rates, and have a preventive consultation with a primary care physician shortly after signing the contract. The assumption is that the doctor will catch any issues in the early stages and prevent them from escalating to peritonitis. At least for this year—and after that, either the insurance provider will change (which won't be Curative's problem anymore), or it's back to the primary care physician.

The company was launched 5 years ago. In December, it raised $152 million in funding at a valuation of over a billion. For 2026, the startup projects revenue of $570 million—which is roughly on par with Lemonade, the trendiest insurance company of the previous generation, which went public in 2020, the very same year Curative was born.

Company website ↗

My beauty and I remembered what we're alive for, jumped in the car, and drove off to France to let the dog run around in the snow. Honestly, I haven't seen any in a while either. Three years, probably.) We'll get to Chamonix today and be back in London next Saturday :)

For anyone who doesn't know, this is what the train under the English Channel looks like.) You drive your car onto it and chill for half an hour. Boom, and you drive off somewhere completely different :***

A strange fact I've been thinking about for a long time: this tunnel was first designed in 1802, before electricity or any kind of engine had been invented (?!?!?!?!). They just decided to dig a tunnel under the fucking sea by hand.)) And they designed two levels right away, with a horse-changing station halfway through, lit by oil lamps. But they only succeeded once some proper tools came along, in 1987. Or rather, they started digging in 1987 and finished in 1990. Thirty reinforced miles under the sea in three years. Not bad either.)

But 1802... Can you imagine the people who lived around here back then? The Elon Musks of their day, probably?)

Damn, I hadn't used Agent Mode in GPT since it appeared. It seemed to get worse almost immediately and looked completely useless.

But today's experiments showed that Agent Mode is a gazillion times fucking better than Deep Research.

For example, today it went through hundreds of vendors selling microcement in London in a couple of minutes. I don't understand how it does that so fast. It gave me all the names I needed and the price range from bottom-end brands to ultra-luxury. It explained how they differ, what people say on Reddit, and put everything into a presentation for me too.

The presentation looks pretty shitty, but the content is incredibly good.

I ended up talking to the builder in a way that left him wondering where the hell I'd learned all this, when we'd met two hours earlier and I couldn't even pronounce the word.

Agent Mode can even add things to shopping carts and try promo codes, picking the site with the best combination of product, discount, and total cost with delivery. Things like that, which are genuinely exhausting to do manually.

Vibe coding isn't necessarily about a bot that automates something. It can simply be an interface you find much clearer and more pleasant for something you do every day.

Nothing in my chess project automates anything. In the business system, it's the opposite: lots of things are automated because manually checking or doing the same thing over and over is fucking exhausting.

There aren't any automations in the donation tool for YouTube streams either. I simply made the interface I wanted for that function and brought the payment options I wanted into one interface.

I also think features and automations are like an appetite: the more you try, the more you want. I didn't plan the chess bot at all. First I got a board and pieces that two people could move. Then I just kept thinking, “So what now?” “What's next?”

Lots of things turned out to be shit, and I threw them straight out. Some things I thought would be shit turned out to be pretty good. They led me to the next two or three features, and so on.

I'm thinking of making myself a proper call scheduler, for example. I've been sick of Calendly for ages, and Zoom Scheduler just broke, besides being crap too. I want a scheduler that works exactly the way I need it to. So I don't spend 15 minutes every time trying to find how to change my availability among 500 buttons whose purpose or audience I don't fucking understand.

You can eventually turn pretty much anything into a product for the market. But if it doesn't even work for you, the market certainly won't fucking need it.

I think this is where most startup founders fall apart. They make who-knows-what for who-knows-whom. Something they don't even intend to use themselves. Just because they imagined something or heard something somewhere.

The result is a product that sounds like something on paper, but in practice nobody even wants to test it, let alone pay.