Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More

Owned by Ray

Watch me build apps and automate my business with AI. Get the prompts, workflows, and honest breakdowns from a 12-year Apple engineer. Free.

118 contributions to Start My AI
Newest coding models flood repos with test slop
The newest coding models write a lot of unit tests. Most of them restate the code they were written after, so they always pass and catch almost nothing. They run fast, but they break on every refactor and the agent burns time fixing tests instead of the feature. I now test with E2E runs that exercise the real feature and leave an artifact I can check and rerun. If a piece has to be tested alone, the model writes down how it could fail before it writes the code. The prompt I used to clean out mine. "This repo is full of low-signal unit tests. Delete every one that wouldn't catch a real bug our E2E tests miss. Fan the work out across parallel subagents, and use pstack and poteto-mode. Then add these rules to AGENTS.md so future agents stop writing them. Never write unit tests after you write code. Highly prefer E2E tests as the sole testing mechanism. Use them to verify complex features work. At the end of E2E tests, produce a verifiable and repeatable artifact. If you must test a system in isolation, first write down all the ways it could fail, then write the code."
Let's talk computer use
I tried the codex app's computer use functionality last week with 5.6 luna. I had it delete an email, then delete an email based on the sender's name. Then change my brave browser's default search engine from brave to google search. It oneshotted everything. I was surprised because a year ago I tried something similar with computer use and gpt-4o. It failed miserably. Codex only released in feb and it's computer use is already so good. I am in the construction industry doing a startup and computer use has so many use cases, it can input hundreds of items that have to be normally written manually. Saves hours of time. Haven't been this impressed about something ai related in a while 😁
0 likes • 23d
@Sabih Sarowar it literally controlled Blender and helped make a ton of 3D assets for my game. I'm currently having it make a bunch of icons for my onboarding for my app as well
1 like • 12d
@Maria Martins I’ll be cooking when I get back from vacation after Oct 5th
Out for members now. Public at 5pm. I Made This So My Friend Could Follow the Conversation
New daily vlog is out for YouTube members right now (Let Me Cook and higher), and it goes public at 5pm PT. I Made This So My Friend Could Follow the Conversation. If you are a YouTube member too, you can watch it now. For everyone else, 5pm. https://youtu.be/SRHxCTnxQ-Y
Out for members now. Public at noon. My First Day With Cursor Projects
New daily vlog is out for YouTube members right now, and it goes public at noon PT. My First Day With Cursor Projects. If you are a YouTube member too, you can watch it now. For everyone else, noon. https://youtu.be/aj9LNTMtCo0
4
0
Grok Bot with T3 Code
I’ve been running pstack inside Cursor for a while. It's a great setup with one problem. Every delegate runs at Cursor usage rates, so the models I actually want on hard tasks were too expensive to use by default and maxing out my free other usage that I prefer go to bug bot. Meanwhile I'm already paying for Claude Max, ChatGPT Pro, and SuperGrok Heavy. I wanted the pstack playbooks with my subscriptions doing the work. So I forked it. Grok Bot became the orchestrator and T3 Code became the execution layer which would give me similar workflows to using grok bot with cursor cloud agents except the cloud is my hardware, and I can use my subscriptions instead of cursors api rates for my other models as well as incorporating my local models and open router if necessary. And I can do it all from my phone which allows me to leave my office and still get work done :) Grok Bot holds the router, the playbooks, and the table that says which model handles which kind of job. It decides what runs where and writes the brief. T3 Code runs on a Mac Studio in my home network and wraps the coding CLIs I already pay for: Claude Code, Codex, Grok Build, and a local model. For each step, Grok Bot opens a T3 thread on the right provider, sends the brief, waits, and reads the results. My original plan was to have Grok Bot use T3 Code the way I do: through the app. That went badly. T3 is built for a human with a screen, and the Bot struggled to drive it reliably. It could get a thread open, but sending work in, knowing when the delegate was actually finished, and pulling the result back out was fragile every time. An orchestrator that can't tell "done" from "still thinking" isn't an orchestrator. The fix was to stop asking the Bot to use a UI and give it a tool shaped for a bot. We built a small command-line tool with exactly the handful of actions the playbooks need: start a thread, send it work, wait for it to finish, read what came back, cancel it. Every action has a clean start and a clean end. Once the Bot had that, the pilot lanes ran hands-off. That's the biggest lesson in the whole project: if an agent is fumbling a tool, don't write a better prompt, build a better interface.
1 like • 24d
Thank you for sharing this workflow and please let me know when you do have the GitHub link. I'd love to check that out. I think this is a really great alternative because there are a lot of other people who are building outside of Cursor and this sounds amazing. Great job figuring it all out!
1-10 of 118
Ray Fernando
6
1,492 points to level up
@ray-fernando-4666
12y ex-Apple. Live AI coding streams Tue/Thu 10am PT on YouTube. I break it down so you build with confidence.

Active 1d ago
Joined Apr 21, 2026
Los Altos, CA
Powered by