Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More
Clief Notes
49.3k
5.0
Free
5 contributions to Clief Notes
My tiny mailbox for Claude and Codex handoff automation
Files carry the words The next agent reads the task No more copy-paste I’m a novice at building software and learning as I go. I run a CNC workshop, and on a long project I had Claude Code doing the work and Codex independently reviewing it. The reviews were useful, but I was copying every handoff between them. So I built Looper. One agent leaves a request in a small folder of readable files. A second agent answers it. If that run fails, the request stays waiting. A fresh session can read the folder and pick up from there. It started with code reviews, but I’ve since used the same handoff for testing, research and getting a second agent to check things the worker couldn’t. Claude and Codex have taken both roles. I’ve also run a local Qwen model through OpenCode and Ollama for a review. That local route has only had limited testing, but it showed me the point: the files hold the job, so Looper isn’t tied to one model or agent runtime. The approach came from what I’ve learned about ICM here, especially the work of Jake Van Clief and David McDermott. The sessions can come and go; the folder keeps what they need to continue. And yes, Looper is partly a nod to the film. During development I jokingly called it Bruce. A job arrives, someone deals with it, and eventually the loop closes. I’ve found it genuinely useful and have made it public and MIT licensed. If you’re carrying messages between agents, or want another agent to check work without losing track of what’s been reviewed, give it a try. It works with plain files, and there are optional helpers to automate the handoffs. https://github.com/shedstalker/looper I’d love to hear what you use it for, what breaks, and what you’d simplify.
1 like • 2d
@Leo Saraiva Good point, Leo. In Looper, “waiting” means it hasn’t been answered yet, even if someone’s working on it. The headless driver locks the reviewer role, but two manually started reviewers could overlap. Your "claimed/" idea makes sense if I add multiple reviewers. Thanks!
1 like • 2d
@Patrice Roatan Quebecois Thanks Patrice! That’s exactly it. I tried to make it as small as possible: leave the request there until it gets an answer, and let me swap models when I need to. Glad that came through!
Agentic AI Governance - two sources I found interesting
I work in AI Governance and am currently doing a course on the EU AI Act, NIST, ISO 42001 etc. while building AI workflows with the ICM method. What I realize is the huge gap in Agentic AI governance. With narrow AI it was okay, GenAI made things more complicated, and with Agentic AI it's become a nightmare, from a pure governance perspective. Here are two papers I found interesting though. 1. OWASP – Top 10 for Agentic Applications OWASP's answer to agentic risk is essentially a design principle, not a control checklist: give an agent the minimum tools, permissions, and autonomy the task actually requires. They call it "least agency," analogous to least privilege in classic security. I think that most agentic failures aren't the model being wrong, they're the model being wrong while empowered to act. So you govern the action surface (which tools, which scopes, which actions need a human sign-off), not just the model output. That's something you can audit, which makes it governable. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ 2. Sinha et al. (2026) – Human-on-the-Loop Orchestration for AI-Assisted Legal Discovery (arXiv) The complementary piece: once an agent does act across multiple steps, errors compound silently and they call it "trajectory collapse." An early misstep propagates through every downstream decision, and endpoint metrics (precision/recall) never see it. Their answer is uncertainty-gated escalation: the agent hands over to a human only when calibrated uncertainty crosses a threshold. In their (synthetic, preliminary) simulation, routing under a quarter of cases to human review cut the critical error rate by ~61% vs. full autonomy. This means that "human oversight" stops being a checkbox and becomes a measurable dial, a threshold vs. workload vs. residual risk. https://arxiv.org/pdf/2606.19812
0 likes • Aug 13
@Mira Bradshaw I keep thinking about this like my workshop. A CNC machine can do incredible work, but you still need the right drawing, current revision, tooling, job order and someone responsible for checking it. None of that is new. It’s just basic production control. My first AI setup was very free-form and honestly worked surprisingly well. But I only wanted help quoting jobs, and it gradually showed me how much basic structure I was missing. That led to spending far too much time building my own governance around it. The upside is that once those foundations were there, every new skill or action became much simpler because it didn’t need to carry all the rules itself. So how far back do I need to go? Are we really inventing AI governance, or relearning old business, production and records principles that got forgotten while everyone chased the shiny parts? My favorite "rule" I wrote was: keep governance proportionate to the risk and drop any ceremony that adds no real control. I'm not one for fluff.
0 likes • Aug 14
@Mitanní Spruill-LeSueur At the moment, mainly human review. At the moment I mostly catch it as a friction point while HARRIS is doing a real task. Something feels wrong, so I stop and trace which logic it used. That gets returned to development rather than HARRIS changing its own rules during the live job. The fix is patched and tested separately. If the same problem comes back after I know it was patched, that tells me the rule exists but wasn’t loaded or applied. HARRIS can flag some missing information and uncertainty itself, but human review is still important during the alpha. The limitation I’ve hit is that HARRIS currently sits on top of ChatGPT as rules and context. I can build those rules as carefully as I like, but ChatGPT can still fail to load or follow one. That’s what happened with the quote. I’m now trying to put HARRIS underneath the LLM instead. The system would load the required information, check the steps, stop unapproved actions and verify what actually changed. The LLM could still reason, but it wouldn’t get to decide whether the governance applies. Have you seen a lightweight way of doing that without building an entire enterprise platform?
The Condom for Claude Dilemma
Not sure who's seen the meme. I found it pretty funny for a second until I thought about how true it actually is. After being let go from my last job as an SRE (Site Reliability Engineer) with a pretty good severance package I decided that I'm in no rush to get back onto the bandwagon. I spent the next 9 month doing pretty much anything that is not working. Had a double hernia surgery, slowly recovered, went out a lot, traveled, spent a decent amount of time snowboarding and in the mountains. The last 4 months I traveled through South-East Asia. Started interviewing for a position while I was in Vietnam and was hired as a Cloud Engineer for a Cyber Security firm in the beginning of the month. I'm saying all that to highlight my lengthy disconnection from the state of affairs when it comes to using AI in the workplace. My current employer expects lightning speed in everything. Something that usually took at least a week to write, test and validate properly, including all the quality gates for review and deployment is now expected to be done till end of business day. This leads to an interesting conundrum. I find that people have no idea what they are doing and delivering and have no time to even try to figure it out. Junior people are writing code that they don't understand using Claude, but with speed 10 times faster than any Senior Engineer I've ever met. They immediately push their code and the more Senior engineer that is supposed to review their work is pretty much doing the same thing as he is pressed for time like their is no tomorrow. They ask Claude to review the code and immediately push the changes up the chain. So it's basically Claude writing code that later is further reviewed by Claude, fixed by Claude and pushed by Claude to production. In this scenario I don't think we're even a condom for Claude anymore. If we are, we are definitely a broken one. There is no actual review, no actual quality standards or gates. What's left is an exhausting sprint that leaves everyone worse off than when they started. Maybe we maximized shareholder value at least. What's striking to me is that this is the state of affairs in a security provider. Can't imagine what's happening in less rigid companies/environments.
3 likes • Aug 13
@Scott Smith I think this is where governance changes the value of the reasoning. If Claude writes the code, reviews it, fixes it and then approves it, the model is reasoning in a circle with no outside reference. It’s like my CNC making its own drawing, cutting the part, checking it against that same drawing and deciding the job is correct. It becomes a snake eating its own tail. Without an independent standard or real check, wouldn’t the code slowly drift and degrade?
1 like • Aug 13
@Scott Smith Time to touch grass for me too.
Should the agent be allowed to propose changes to relationships, or only to values?
my question in 1 sentence : Should an ICM agent be allowed to propose changes to relationships (fields pointing one record at another, like ownership) or only to scalar values (self-contained data such as amounts, dates, statuses and free text) — and is that distinction even a legitimate permission boundary, given that Jake's model draws the line on role, route and sensitivity, never on field type? ------------------------- The context I run an ICM where the catalog is generated from the system that holds the actual records, and where the agent can propose changes back through a human gate. Reads flow one way, proposals flow the other, and nothing is written without me approving it. When I built the proposal path, I made a decision I now think was wrong. I allowed the agent to propose changes to values — amounts, statuses, dates, free text — but I blocked it from proposing changes to relationships: the fields that point one record at another. My reasoning at the time was that a pointer is structural, and structure should be changed deliberately, through the interface, not through text. Re-reading the ICM Architect material, I no longer think that distinction holds. Jake's permission model is built on role ("agents inherit the access of the role they act for, never more"), on route ("the permission lives in the routing, at the data layer or along the route"), and on sensitivity ("shelve by sensitivity"). Never on the technical type of a field. And relationships are not second-class in his model. Invariant 8: "Links make it a graph… the edges are what let an AI move through your work the way you do." His own node template carries owner: <accountable person or role> right next to consumes: and produces: wikilinks — all of it ordinary frontmatter. So I may have invented a governance rule that is really just an implementation constraint wearing a costume. --- The concrete example The field is ownership — who is accountable for a record. Today: Me : "Who is accountable for this one?"
1 like • Aug 13
@Ramon Tilanus I gave mine a simple rule: it can suggest something outside its authority, but it can’t act outside its authority. In my workshop, anyone can suggest moving a job to someone else, but changing who is responsible for the final sign-off needs my approval. So I’d let it propose the change, then make the approval match the consequence. I think that’s where governance becomes part of the context. The LLM is given a task, the information it needs and the boundaries to work within. It doesn’t make the whole decision on its own. Am I thinking about that the right way?
Grandma passed and my uncle has cancer
Hey y'all, not a sympathy post. This isn't me saying I'm taking time off for anything. In fact it's the opposite. My grandma passed away last week and I was cleaning out her apartment with my mother and talking to my aunt (both of them it was their mother). I can tell they're struggling. It's why I work so hard honestly on the software community on everything else. Almost none of the money goes to me. I give it to my family, my wife my friends. It's what makes me happy. I'm building a company specifically so my family doesn't have to struggle. Been helping out where I can but obviously been focusing a lot on the business. My uncle also has cancer and so my aunt has been taking care of him and also at the same time dealing with a death of her mother . Best way I can contribute has been through cash and some love here are there. With that though means I don't always have the time to do what I need to do. So honestly, I'm asking for your help if you're willing to donate to my aunt's GoFundMe on top of the help that I'm giving. I've just donated a bit and would love other people to help out. To me, if I work hard I can make sure that my whole family is taken care of. But in the small ways I can help where I can I will do my best too And that goes with asking for help from others. So if you all could help me out help them out. That would be amazing https://gofund.me/45b6655ba
2 likes • Aug 10
@Jake Van Clief Tough times - hang in there and trust your gut.
1-5 of 5
Nicholas Atkins
3
36 points to level up
@nicholas-atkins-9464
Melbourne boatbuilder, designer and CNC business owner focused on quality, smart systems and practical problem-solving.

Active 14h ago
Joined Aug 8, 2026
Melbounre
Powered by