AI agents
Company context is infrastructure
We brought our code, tools, and team knowledge together, so that our agents could work smarter and we could work faster. Here’s how.
In the first few months of 2026, we discovered that with each new connection, our agents became more capable.
Connect one to Slack and it could piece together why we'd built something a certain way. Give it access to Notion and it could find the documentation. Let it read the code, and it could trace a customer's problem to the exact bug responsible.
We kept connecting things. Eventually, that pulled support, design, growth, and engineering into the same working environment: our codebase.
Experimenting with agent context
Going into 2026, the agent world felt a bit like the Wild West. We were trading ideas about what these tools could do and trying them out ourselves.
One of those experiments became Replee, an internal Slack bot we initially built to investigate Sentry issues. It had access to Slack, GitHub, and Linear, and we quickly found reasons to make it more general-purpose.
We also started making our internal tooling accessible through an MCP server. The bot could call those tools to look up product information instead of implementing each lookup itself.
Alongside the coding agents engineers were already using, we began bringing the rest of the company into the loop. In February, we organized an “agent learning” party for non-engineering teammates on how the bot worked, with reading from Ramp and Stripe and a request to come with questions. Support learned how to use the bot to find information and investigate issues. Our content lead and our design lead, who also help drive product direction, started using agents to streamline their day-to-day.
In the weeks that followed, support was using the bot on real tickets. It could look up product information and consult the documentation and code, but first it needed the customer's conversation. Our support QA workflow relied on tickets being forwarded from Pylon, our support platform, into Slack, and that relay was missing tickets.
In March, we gave the bot direct access to Pylon's conversation API so it could read the messages at the source, without depending on what had made it into Slack. That access also let support point it at a ticket and get an investigation back as an internal note in Pylon.
Replee accumulated integrations and instructions as people found more uses for it. Most of what made it useful was everything we'd taught it about how we worked as a team.
These experiments were useful, but the context was still fragmented. The bot was its own project, our internal tooling was still taking shape in a separate repository, and the product code lived elsewhere.
Agents could interact with these pieces, but working across them meant moving between projects and reconstructing context between tasks. As the questions crossed more of those boundaries, it made sense to bring the tools closer to the code they worked with.
Bringing code, tools, and knowledge together
In March, we moved our internal tooling and its existing MCP server into the monorepo. The pressure was fairly practical: people outside engineering were running parts of the product locally to get work done, while our internal tools lived in a separate project that was awkward to share and edit.
We wanted tools the whole team could use, and agents that could see the product code while working on them. We brought the existing app and MCP into the repo in stages.
With this change, an agent could use a tool and then read the code that implemented it. If the result was confusing, it could inspect how the tool worked. If we needed to change the tool, the product code it interacted with was there too.
We also established .agents/skills as the shared home for agent guidance, so OpenCode, Cursor, Codex, and Claude Code could all use the same instructions. Scripts generated symlinks into the directories each harness expected, and CI enforced that shared-source pattern. Each skill had one source file to update, without anyone having to remember which other copies needed the same edit.
As more guidance moved into the repo, people could work further beyond their own area. Support could trace a question into product behavior; engineers could use design and copy guidance while building a feature.
The agent could follow those linkages without us mapping out every project first. I could ask about an unfamiliar project, get a lay of the land, and start contributing much faster.
But the code couldn't tell an agent everything. Where to find logs, how to investigate a session, and what to check before shipping came from the engineers doing that work. Other teammates brought the same practical knowledge to our content voice, design standards, and support procedures.
Turning specialized expertise into a team resource
Support's feedback changed what an investigation needed to return. When the bot returned technical findings, they asked it to include the relevant knowledge-base articles too. That gave them an explanation they could check and documentation they could send to the customer.
We turned repeated support work into instructions the agent could follow: which sources to consult, what to investigate, and how to leave findings for the team. A ticket tag lets support start that workflow without writing a prompt from scratch. We also put together a runbook so the rest of the team could learn how to use it.
Our content lead added our marketing brand voice as a shared resource: it defined how we should sound in newsletters, product copy, and other writing, with channel-specific guidance and examples. After that, she did the same for in-product copy and our own agent’s voice when conversing with users in chat.
Soon after, she was using it on a newsletter and refining the guidance, such as establishing voice conventions and writer sign-offs. In May, she had the skill copied into our product repo so the rest of the team could use it. The expertise came from her; the skill gave other people's agents the ability to apply it on any surface.
Our design lead helped shape the design guidance around the workflows she actually used and our design system. She also flagged details agents kept missing, like compatibility with a project's theme. Those conversations gave us specific things to ask agents to account for when building new UI.
These were people putting their own expertise into something the rest of us could use. People could iterate on product and design without making our design lead pull their hair out, or spin up in-product copy and in-app announcements without consulting our content lead.
We even got engineers writing decent copy.
Each of those improvements gave the next person a better starting point. But as more people contributed their own workflows, the skill catalog ballooned faster than our system for organizing it.
Optimizing an expanding skill catalog
As people rode the high of accelerating workflows, more and more skills were added to the repository. For a while, adding another set of instructions usually made the agents more useful, so we kept doing it. With the models and harnesses we were using, explicit workflow guidance made enough of a difference that the extra context was usually worth it.
The number of skills we had skyrocketed and we often found ourselves with several skills trying to do the same thing.
Our first big cleanup in May was about finding the common structure in those workflows. When we asked our design lead which design skills she relied on, she mainly named three. She hadn't realized we had 27.
Several skills described similar jobs, leaving agents to choose between overlapping instructions. Related review, debugging, coding, and design workflows didn't all need separate entry points.
We grouped them behind broader skills, keeping the specialized guidance in files an agent could read when the task called for it. The catalog went from 92 skills to 33, and the measured description text fell from roughly 4,200 tokens to roughly 2,000.
Reducing the catalog only helped if agents could still find the guidance people depended on. Before consolidating skills, we asked their users which workflows mattered, then asked them to flag any drop in quality as they kept working. My version of that request was: "Obviously be loud if this makes your agents stupid."
After that cleanup, we kept tightening our skill for writing skills. Drawing on Anthropic's guidance, open-source examples like Superpowers, and what we were learning from our own agents, we documented how to justify a new skill, check for existing guidance, and organize its instructions.
That work grew to include testing discovery with fresh subagents. We gave them realistic requests without naming the skill, then checked whether they found the right instructions. We checked which skill they opened, which reference files they followed, and whether their next steps matched the workflow.
We also tested nearby requests that should route somewhere else. These agents hadn't helped write the skill, so we could test whether a normal request was enough for them to find it.
The May cleanup gave us a better way to organize the workflows we already had. As the product and our use of agents grew, though, the catalog needed to cover more kinds of work. By late July, we had more than 70 top-level skills again.
We were also bringing recurring work into the same system. In July, we moved 14 scheduled workflows into the repo so their instructions could be maintained as skills. Initially, many had their own nested SKILL.md files, which pushed the total file count up further.
That exposed a different limit. Agents first see skill names and descriptions, then load the detailed instructions they need.
Even when the workflows were useful, their combined discovery text could get too large. Some descriptions stopped making it into the context.
One particularly ridiculous casualty was our writing-skills skill. We'd put all that work into teaching agents how to justify, write, and test new skills, but thanks to W's unfortunate position in the alphabet, its description was getting cut off in some harnesses. Sloppy skills were being added without going through the checks we'd built for them.
Our skill for writing skills had gotten lost among the skills. I gave it a name starting with A to move it toward the front of the alphabet.
The more durable fix was to make the higher-level workflows do more of the routing. A review request could load one review skill, which would then direct the agent to more specific guidance based on the code and the task. Specialized instructions stayed available underneath it, without every detail competing for attention at the start of every session.
That's the useful part of progressive disclosure: the request determines which context the agent reads next.
Our automations are a useful example of how this works. One entry point directs the agent to the instructions for a specific workflow, which can then link to any supporting data it needs. All of those jobs stay accessible without each one needing its own entry in the initial catalog.
The directory kept one SKILL.md for routing, with individual workflow files underneath it:
.agents/skills/automations/
├── SKILL.md
├── customer-frustration-digest.md
├── ...
├── pokemon-cycle-names.md
├── product-updates-reminder.md
├── publishing-health-weekly.md
└── ...With those routing changes in place, we tested discovery again with fresh agents to check that they could still find the workflows they needed. We inspected the catalog they actually received, shortened descriptions, and set explicit budgets to keep the entry points within the available space. Mechanical checks caught invalid or oversized descriptions.
The later cleanup brought us from that peak of 119 skill files to 42, with specialized workflows kept as ordinary reference files behind those entry points. The measured description text fell from over 10,000 tokens to under 2,000.
These cleanups changed what we considered a finished skill. The instructions had to be useful, discoverable, and worth the context they occupied. Writing them down was only the first step.
By September, the catalog had grown again to 68 skills, but the measured description text was still around 3,150 tokens and within our budget. We could keep adding useful context without advertising all of it at the start of every session. That mattered as the repo took on more than development instructions.
One source of truth for knowledge and workflows
While we were figuring out how to organize that context, we were also changing which agents used it.
By May, we were finding that Cursor, Codex, and Claude Code could handle much of the work we'd built Replee for, especially when they ran with our codebase's context. As my boss put it in Slack: "agents are commoditized, it was a good 3 months." We could put more of our effort into the tools and company knowledge those agents needed, without maintaining our own bot around all of it.
That transition wasn't just a matter of picking a different agent. Support still relied on Replee's access to customer conversations and internal tools, and those capabilities needed to come with us.
Debugging got much faster as we connected agents to more of our systems. With access to Grafana, Sentry, PostHog, and the codebase, an agent could follow a customer's session across services, match timestamps and IDs, and trace an error back to the code responsible.
Each integration gave it another part of the investigation it could work through without someone copying information between tools.
But those capabilities initially depended on what each person had configured. Two teammates working in the same repo could have very different results from their agents.
In June, we started connecting agents through Executor, a shared gateway for our tools. That made those connections available across the team, while our debugging skills taught agents how to use them together. Adding an integration or improving an investigation workflow could then help everyone's next investigation.
Alongside that work, the repository was becoming a home for more of the work around the product. In June, customer documentation moved in from a separate repo, along with the source of copy and workflows for our marketing emails. The content, the instructions for working on it, and the code it referred to could be maintained together.
Our growth lead brought email and sales work into the repo. Our demo-request emails and their follow-up playbooks moved out of HubSpot's editor, so updating them now means asking an agent (accessible from our Slack) and syncing the change back.
The same thing happened with our internal CRM, which has been used to run 1-1 email campaigns as part of targeted user positioning experiments.
With content and workflows in the repo, scheduled jobs could use the same instructions as someone working interactively. The scheduler says when to run; a repo-owned skill says what to do.
Whether we start a workflow from a local Codex session, a Cursor automation, or another agent environment, it can read the same instructions. Moving the job still requires the right tools and access, but it doesn't require rebuilding its instructions from a prompt hidden in someone's account.
One of our weekly Cursor automations gives our Linear cycles Pokémon names in Pokédex order. This is its entire prompt:
Read and follow .agents/skills/automations/pokemon-cycle-names.md exactly. Do not improvise beyond what the skill specifies.When we moved this job over, the existing Linear MCP integration didn't expose the cycle-renaming operation available in Linear's API. We added that operation to our internal MCP server in the same PR as the Pokédex and the instructions for using it. The tool implementation, workflow, and supporting data could all be reviewed together.
That's the same benefit we'd started getting when we moved our internal tooling into the repo, now applied to the work running through it. If a run goes wrong, we can inspect the instructions and the code behind its tools in one place, make a correction, and test the workflow ourselves. The next scheduled run and the next teammate asking an agent to do that work can use the same updated instructions.
By then, those shared workflows were also changing how we handed work to each other.
A shared system we constantly improve
By this point, a short Slack message could draw on months of work we'd put into shared context and history. Someone reporting a problem didn't need to know which service owned it, which logging tool to open, or which skill to ask for.
We started treating an agent investigation as part of reporting an issue. In July, we were explicitly asking people to involve an agent when flagging a problem to a teammate, so an investigation could already be underway.
Someone in support, design, product, or even growth could kick off the investigation by tagging an agent in a Slack thread. The agent could start with the conversation already there and follow the investigation workflows in the repo before an engineer picked it up. A Sentry alert could enter triage automatically, and a teammate could ask for a deeper investigation in the same thread.
An engineer could arrive to find relevant logs, a diagnosis, or a proposed PR waiting for review. The person responsible still owned the outcome, but getting started no longer depended on having all that knowledge in your own head.
The path from an issue to a reviewed fix was getting shorter, and more people could move work forward. Over the same period, the repo was getting busier. These are monthly merged PRs:
By late July, support workflows were moving into the monorepo too, where they could use the same debugging and investigation guidance.
In August, our support lead investigated a billing issue with an agent: customers with an overdue invoice were getting an error when they tried to update their payment method. The investigation traced it to a subscription lookup that excluded overdue subscriptions, blocking the very people trying to fix a failed payment.
By the time support brought the issue to engineering, they already had a diagnosis and a PR to fix it. Twenty minutes later, after engineering review and deployment, they'd verified the fix in the customer's account.
The customer report supplied the problem; the repo supplied the code and context to work on it. The person starting the investigation could describe what was wrong without already knowing where the fix belonged.
When our emails encountered deliverability issues, our growth lead was able to pinpoint the exact issue down to a DNS record and HubSpot setting with the help of an agent and the full history of our past email sends. The results of the investigation were written as deliverability rules into our marketing email skill for future reference so that the same issue would never happen unnoticed again.
Our content lead built workflows for lifecycle emails and weekly content. Onboarding emails and interrupt emails were now set up to be triggered by specific user-tags and pre-set time stalls.
Workflows were built to research, draft, and review ideas for our social and longform content. Once a week, one job researches SEO topics and drafts blog and newsletter ideas; another scrapes social topics from our Slack, code commits, and meeting recordings, then checks them against our voice rules.
Voice, diction, and writing conventions rules are enforced by a mix of code and automated checks. Nothing publishes until someone approves it, the jobs just make sure there's a draft waiting for her to review first thing in the morning.
Whether a story is worth telling or a product experience is optimal still requires human judgment, but the rules we can check no longer depend on an agent remembering every instruction. Instead, they come with full history on our product and users, sometimes with even more context than what we retain ourselves.
Design feedback can become instructions for the next feature, and support's experience and tickets can improve the next investigation. People can improve the system from the part of the company they understand best.
Today, everyone at Replo uses AI day in and day out. Each of us brings unique expertise to the same shared working environment.
A fix to our debugging instructions can help support handle the next ticket and help an engineer investigate the next incident. Neither has to know who first ran into the problem, or track that person down to learn what worked. It’s already preserved in our collective memory, a shared knowledge base that all our tools, code, and agents have access to and can contribute to. That's the part I find most exciting and innovative about where we've ended up.
We're still exploring what else becomes possible when everyone can build on each other's work and we each have just as much context as the next person. Whoever joins will help shape that, too.
If that sounds like a place you'd enjoy working, take a look at our careers page. Even if you don't see an exact fit, feel free to say hello.