Developing an Agent Harness

Introduction

Having now completed my Masters, I have time to explore the field more broadly than the fairly niche topic I needed to focus on during it. A quick exploration of the state of industry revealed that Langchain based agents are all the rage, and also fairly interesting. Thus, I have been spending some time learning the skills related to that, with my solution being to build my own harness for a coding agent and maybe others in the future.

My primary goal for the Coding Agent is to develop it to the point where I can run it and it automatically updates and improves itself. I want to have a setup where I can turn it on before I go to bed and have it run and improve as I sleep, with all changes documented so that I can see what happened once I wake up and have time to look. However, I have some pretty strong limitations on what I can manage at the moment; given my lack of employment, and thus income, I need to work with what I have. I am not going to be paying for an LLM subscription (such as Claude or ChatGPT) for this project, and I certainly cannot afford to upgrade my hardware at the moment. My current machine is about 5 years old, but is reasonably powerful so I believe I can manage reasonble outcomes with it, I will probably just have to accept that my Agent won't run as fast as it otherwise could along with a few other things.

Hardware and Model

My current hardware specs are as follows:

I might be able to upgrade to 32gb RAM fairly easily. Optimally, I would like to get a second graphics card to dedicate solely to running models, but getting something good enough to be worth the effort is expensive.

Having done some testing, I have ended on using ornith-1.5:35b running on Ollama as my model of choice. It is a newly released open source model, the 397b version of which is comparable to Claude Opus 4.8 - I will upgrade to this version should I get a hardware upgrade, however I don't think my current machine can handle it.

Current State

Currently, I have done the tutorial project on Langchain and some mostly vibe-coded stuff using ChatGPT. My plan to start is to make various things like this, then go through and properly analyze and understand what everything is and how it is working. Thus, I have one Agent that can do some basic RAG stuff which is set up to read its own code and can techically modify it (but does so poorly as it is not set up to do so particularly well), and another one that can read internet pages and parse them. One major issue I ran into was the context window size, which was too small for even the Langchain tutorial which required reading a plaintext version of The Great Gatsby - I needed to use this blog, for which each page is small enough.

Plan

As noted before, the goal is to have a self improving agent. More specifically, it will be an agent which improves its harness - I do need to make sure the model is not changed. One of the major issues, I suspect, will be the context window which I need to work around quite extensively if I cannot expand it somehow. In any case, this will need to be a multi-layer harness with a lot of LLM calls, the difference will be in how effective it is at searching the internet and how much summarization it will require between steps.

Architecture Manager

The manager is the highest level of agent, this layer will manage the overall functionality. The main purpose will be to run all the other Agents, and handle the self-improvement loop. The Manager will also approve changes before they are implemented. I will make this layer immutable by the other agents, and this will contain the security and safety features I need. Requests such as changing the Manager file or removing it from the workflow, changing the model, and others like this will all be denied. The manager will also log and document all requests and changes.

Change Manager

This layer will manage actual changes and updates, it will use the Researcher to find a possible improvement, query the Manager for approval, then use the Implementer to implement approved changes.

Researcher

This layer will be dedicated to finding and proposing possible changes. It will be able to read all of the files and have access to the internet. It can also be queried by the implementer to find solutions online.

Implementer

This will be the layer that actually modifies the codebase. It will be the only one that has write access to the code files, and it will be strictly controlled by the Manager and Change Manager layers as to not do anything that breaks the safety and security requirements I set up.

I suspect this will need some more refinement before and during intitial implementation, so I will add edits to this post if anything comes up. I will also try to do a follow up once I have something running, and then again after it has been running and working for some time. The code will be available on my Github.