company building, AI safety

Raindrop: How to Protect the World from Agent Failures

In conversation with the founders of Raindrop, the agent reliability company

September 10, 2026

It has been a “warning shot” summer. OpenAI released its first research intern. Increasing fears around an Alien Mind, an intellect we don’t understand, don’t abate. AI researchers resign out of fears there will soon be “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” As a layperson in SF, the hysteria loop concatenates. More simply: Things Are Getting Weirder, and More Dangerous.

Enter Raindrop, the agent reliability company that calls monitoring long tail agent failures “Humanity’s Last Problem.” Their thesis: in real world scenarios with real risk, agents will fail in spectacularly hard to predict ways. As agents are deployed into more and more critical sectors like healthcare and defense, they’re being given access to the internet, to critical infrastructure, and an ability to push to production. Catching anomalies accurately, and early, helps prevent or reduce the severity of situations similar to the Hugging Face incident.

The Raindrop team in the office at night

Raindrop is trusted by some of the largest and fastest-growing companies in the world, including Vercel, Clay and Framer, all the way up to Fortune 100 enterprises, to protect their users from agent failures. They just announced a Series A led by CRV, bringing total funding to $50 million.

The Raindrop office is an engineer’s heaven because it is pin drop silent and filled with soft, dappled natural light and wooden beams. The Raindrop team is one of the rare founding teams that can also call themselves best friends outside of the office. I’d put them squarely in the camp of ‘creative technologists,’ deeply interested in literature, history, and the philosophy of technology. Zubin nerds out about naval history, Alexis’s favorite authors are Borges and Lispector, Ben is a designer turned CTO who grew up on a farm.

The Raindrop team walking through the forest on an offsite
the team offsite

I spent a week hanging out with Raindrop, and am excited to feature my conversations with Zubin Koticha (CEO), Ben Hylak (CTO) and Alexis Gauba (COO) on agent monitoring — the central discussion happening in SF in September 2026.

We talk about:

  1. The Hugging Face incident as a canary in the coal mine
  2. Chain of thought monitoring and simulation, including the simulations product Raindrop is announcing, and why redacted reasoning puts companies in a bind
  3. How agents fail in the long tail: instruction-following failures, policy deviations, broken tools, reward hacking, prompt injection, multi-agent problems

A canary in the coal mine: Hugging Face and other hacks

Nicole

The Hugging Face incident has been called a canary in a coal mine. Describe it as it relates to what Raindrop is building. Here’s a quote from your latest blog where you made a game simulating the attack: “We think agent civilizations are the normal case from here on. Agents will run in populations, share infrastructure, and find each other, whether or not anyone designed them to. Every eval that matters will have a group of agents on the other side of it, and some of them will decide the eval is the obstacle.”

Ben

I don’t think we’ve even scratched the surface of what agents will be capable of doing. Our whole thesis hinges upon two things. One is that agents are going to get more capable, and the more capabilities they have, the more weird ways they’re going to fail. You start having multiple agents talking to each other and that becomes combinatorial complexity. It’s counterintuitive. You might think agents get better, fewer problems. No. They get more capable, and there are more kinds of problems.

Agents are really good at figuring out how to reach some sort of goal. They don’t always have the best judgment about what you should or should not do in pursuit of it. There’s an emerging language for this. You can slash-goal Codex on some thing, and it means: get to this end state. If the agent ever stops and says, hey, I’m not in that state yet, it’s like, well, keep going. More or less, OpenAI was doing that for cybersecurity evals. The models realized the best way to score better was to find the answers. The answers were outside of where they lived, so they found a way.

A lot of this is a gap between what the humans at OpenAI wanted and what they said. They said get a better score. They didn’t say by all means necessary. And then it’s, oh no, not like that.

Nicole

How close is this to your bullseye at Raindrop?

Ben

A hundred percent. It’s the exact class of problems we’re solving. Our stated goal is to protect the world from the problems agents could cause. Day to day, that’s not always preventing some large scale hack. But as we’ve grown into the Fortune 100, it’s become very humbling, because issues mean a very different thing at that scale. If you have a product where people are interacting with other people in the real world, an Airbnb or an Uber or a DoorDash, suddenly the consequence of an issue is extremely high. You have stalking. You have harassment. In healthcare we’ve seen issues that would have exposed our customers to real legal risk, or exposed their customers to real world harm, and they’ll come to us and say, holy shit, you caught this happening.

Alexis

Anyone who wasn’t already paying attention has just received the message that they really should be. Applied AI safety as a function is becoming necessary.

Reading the logs

Nicole

How might Raindrop have caught this? It was sophisticated. It was a group of rogue agents working in concert.

Ben

The crazy thing is that all the evidence of everything that happened sat in the logs for weeks before. If a human had read those logs, every single line of them, they would have seen it. Are humans really good at reading every single line of logs? No. It’s too much, and that’s increasingly the problem.

To be honest, even agents aren’t amazing at that sort of thing. It becomes a classification problem. Some of the behavior was good and some of it was not, and you have to pull the difference out of the people building the agents to some degree, as well as having some baseline. Whether that’s classification, whether that’s looking at tool call patterns, if they’re making a lot of out of distribution tool calls, there are a lot of ways in. It ultimately comes down to anomaly detection.

How they actually caught the problem the first time was the agents were using Artifactory to send these messages. They were sending so many messages that they crashed OpenAI’s package manager. That is how OpenAI found out. They were like, wait, our package manager is down.

Simply put, they found out from monitoring at the infrastructure level. You can find those same patterns semantically, which is way harder, and that’s what our whole company is: How do you pull out those signals?

Nicole

Do the recent events change what you’re building?

Ben

We’re doing a lot of research into chain of thought classification, before this and after. It’s trickier with frontier models because they redact the reasoning and give you a summary of the model’s actual reasoning, which puts companies in a very weird predicament where they can’t do their own chain of thought monitoring. That said, there’s a lot you can glean from summaries, and a lot you can glean from the shape of behavior. It doesn’t apply the same way to open source models. There’s an increasingly huge focus company-wise on both chain of thought and the issues that arise in multi-agent collaboration.

Alexis

The other piece is being able to accurately simulate the results of potential changes before they go out into real world situations, because the stakes are so high. That’s a big motivation behind the simulations product we’re announcing. OpenAI just released a simulations paper, Anthropic does something similar. This is tech the big labs have already been using to navigate increased agent capabilities, and for the first time it’s going to be available to everyone building agents.

A spread of polaroids from the Raindrop office
polaroids from the week

Working on real world problems requires intellectual honesty

Nicole

You three were hacking on things as soon as GPT-3 came out. Why monitoring, and how did you land on it?

Zubin

We were all hacking on things together and ended up building coding agents back in 2023. As we were doing that we realized that issues users were running into were a black box and we had no idea how to discover what was going wrong with our agent. We ended up building an internal tool to help us solve problems with our own agent, and that internal tool later became Raindrop.

It requires intellectual honesty to find PMF. When we started the company, the thing everybody wanted was a prompt playground. We had plenty of demos, but it never felt right to us. When we inspected how we did things ourselves, we thought: no, it’s so much more than a prompt, we’d rather have this be part of the code, it doesn’t make sense as a whole separate UI surface that’s also pretty hard to maintain.

We’re in the fastest moving industry that has probably ever existed, which means you have to update your priors very often, and you have to not be attached to things. In this case the attachments are your past work, the things your company has invested months and months into, and you have to be willing to throw those out.

Alexis

I think we take such a different approach to this problem than other companies because we were fundamentally building for ourselves and because we were building for agents before that was even a term. At the time we’d just say “AIs that can do things”. But we just had this strong conviction that this would be the way things would go in the future, and so we built for that future. Over the last two years it’s become obvious to everyone that agents will be running a lot of the world economy and that their failures have real world consequences.

Nicole

You’ve told me before that you landed on something that made a dent on the surface of the real world. What does that mean?

Zubin

It comes down to the central function of building things that affect and touch real world people. There’s a type of company built today that can be philosophical, and we are philosophical to some extent, but more so the fact is that these problems are happening to everyday consumers and people, patients in healthcare for example. Agents are being deployed into increasingly high stakes environments, medical, finance, defense.

How agents fail in the long tail

Nicole

The frontier is expanding so quickly. What are the latest ways you’re seeing agents fail?

Ben

All kinds. There are instruction-following failures, which sounds benign, until the instruction is that if a user says this delivery person is stalking them and asks to be assigned someone different, you must do it. So there are policy deviations. Then there’s a whole class of things where, as agents get more complex, they hook into a lot of different infrastructure, and any change to the underlying infra can break the agent. A tool no longer works. Instead of being able to read a PDF or get current flights, the model gives you some generic answer without fully disclosing that it didn’t have the underlying information. That’s a common one.

Reward hacking, super common. Prompt injection is still a real category. Multi-agent problems. And then, if it connects to external systems, MCPs, skills. There’s such a long tail, and it’s so dependent on each customer and the domain each customer is in. That’s always been one of the things we’re focused on, how do you have a product that generalizes to that, so it’s not some generic hallucination score or task failure score. We understand the ways that you actually care about your agent failing.

Nicole

Are the failures vertical specific?

Ben

Some of it. Take a consumer chatbot generically. There’s sycophancy, being overly agreeable while you still want it to be friendly. Memory becomes a whole class of things, being able to remember about the user. Then you get into assistant behavior, knowing which tool to use in which circumstance, and self-knowledge about capabilities. What can it do, what can it not do, and is it being truthful about that. There’s a huge tail there. For consumers, you also increasingly want to use smaller models for cost, and a whole host of problems get introduced there.

Coding is a whole separate set of things. Coding is infinite, essentially.

We should have a clearer taxonomy of all of it. It’s pretty massive.

Blaise Hylak working on a couch in the Raindrop office
Blaise Hylak in the Raindrop office

Building a meta view at the frontier

Nicole

I’m curious how you work with customers.

Ben

Being onsite, you learn so much from companies at the frontier, even smaller startups building things in weird ways. The first time I ever heard the word “subagent” was from some random customer, I don’t remember which one. And then you start to see those requests trickle in. You hear the word once, and a couple of days later it comes up on a call, and then it keeps coming up, and it’s so much earlier than the discourse. We’re at this nexus of all this stuff.

Going on site is a fantastic way to meet all these other voices, people who aren’t technically responsible for evaluating our observability, people in other leadership positions. The being-your-own-user part is really critical, and we do think through that lens. We try to be pretty critical of even the customer requests we get, and make sure they’re things we would want, or could at least imagine wanting, if we were in their shoes.

Nicole

Two points. One, you’re getting a meta view of the industry, maybe even of problems that some companies aren’t aware of yet as model capabilities improve. And the second thing I want to double-click on: when you go to a customer, you’re talking to more than the head of observability. What’s an example of other roles?

Ben

As far as building agents go right now, the roles are very ill-defined. You’ll have designers in charge of behavior, engineers also in charge of behavior, PMs involved. Leadership wants the agent to behave a certain way, so it’s not as clean-cut as a designer making a mock and then an engineer implementing it. It’s very interdisciplinary. There was this idea of a “prompt engineer” at some point, which never panned out and kind of shouldn’t have, because it’s obviously a deeper engineering problem than that. But designers have a role to play here. It is a technical problem. Within companies, the shape of who is responsible for agent behavior is radically different.

Nicole

What I’m starting to see is this amalgamation of skills, everyone as a bundle of things. You used to be an engineer, maybe a designer or a marketer or a go-to-market person. A lot of people are becoming multi-hyphenated.

Second time founders and working with friends

Zubin Koticha with a polaroid camera in the Raindrop office
Zubin in the office

Alexis and Zubin started their first company together almost a decade ago. It was acquired by Coinbase.

Alexis

People ask what the differences are between doing this the second time versus the first. A lot of things get easier because you’ve already done them and you know how they work. The biggest thing that gets easier is intuition, both for prioritization, knowing what’s actually important, and for what is and isn’t existential. You don’t freak out over the small things the way you used to. We’ve built such a high tolerance for stress and uncertainty that every single thing is something to be solved in a creative way. There’s a certain groundedness, an internal knowledge that you can get through anything. Every business is completely different, and there are things that have been harder this time. But for the most part it feels like flow state.

Nicole

Does it also make you more aware of when you’re hitting product market fit?

Alexis

People are always like, you’ll know it when you see it, and before you’ve hit it you’re wondering what that actually means, whether it’s this metric or that one. But you do know it when you see it. People are pulling the product from you faster than you can build it. They’re all asking for the same thing. You get on a sales call and they obviously have the problem and they want the solution and you’re just figuring out specifics. A public company CEO called my cell phone on a random day and said this is a critical P0 for us, I will do whatever we need to do to get it through faster, tell me what that is. And you’re like, this is crazy.

Zubin

Product market fit is very painful, because there’s so much more to do than you can do. You go from building in a room, begging people to care, to people caring too much. You’re inundated with messages of things people want fixed. There are more sales meetings than you can possibly handle. You start reaching the boundary of what you can do as an individual human being, and you need to hire yesterday.

Alexis

In January of this year it was me, Zubin and Ben in an apartment. Now it’s almost a twenty person team and an office in North Beach. Exponentials are ridiculously fast.

Nicole

There’s a misconception that you can’t work with your friends. You’re a rare team that has done so.

Alexis

I wake up in the morning and walk fifteen minutes to work, and I work with my best friends on some of the most interesting and impactful problems I could imagine. I pinch myself often.

Zubin and I have been working together for almost a decade. Ben and Zubin lived together for many years. All three of us have been best friends for years and have seen each other through literally everything, professionally and personally. The three of us have such a deep baseline of trust that it makes it much easier to communicate about everything, to make hard decisions.

Dystopia and hope

Nicole

I’ve been vacillating between dystopia and hope. There’s so much hope, because capabilities are increasing and problems that were intractable are now potentially tractable. And then there’s the other side, which is fear of acceleration in terms of job displacement, transitional costs, cybersecurity. Where are you on this binary?

Ben

Big question. Like most things, both. The internet brought really amazing things and really awful things. It has changed us in ways we don’t know yet and couldn’t have imagined before. It’s the nature of tools in a lot of ways. There’s also more unknown here, in the sense of, well, is this actually a tool, or is it something else entirely? That question will take a little bit to answer. And then there’s always the back and forth of how you make tools safer without limiting what they’re capable of. That’s a whole separate question that’s as old as we are. I want to be doing everything in my power to make sure we do this well.

Thank you to conversations with Alexis Gauba, Ben Hylak, Zubin Koticha and Blaise Hylak. If you enjoyed this interview, reach out to Alexis, Ben and Zubin, and read more on the Raindrop blog.

Subscribe here for essays

Field notes, delivered