Live demo Featured on Ray Fernando

Watch AI break into a live app, then prove every hole.

Ray Fernando (28.8K subscribers) put HackZero on a deliberately vulnerable app, live. With nothing but a URL, the agents attacked from the outside, proved each finding with a replay, and could open the fix as a pull request. Here is the full session, with chapters and the transcript.

Can't see it? Watch on YouTube · published 2026-06-19 on Ray Fernando.

What happens, chapter by chapter

Jump to 2:05

The founders

Ryan Cruz (CEO) spent a decade in cybersecurity CTF competitions, taking first place in Mexico and second in Latin America, and increasingly leaned on AI to compete. Cuau has 14 years building startups through Y Combinator and Techstars across marketplaces, software and hardware, the last two heads-down with agents. Their thesis: agents let you ship fast, and the same speed ships vulnerabilities.

Jump to 5:35

A live pentest, one click

They run a pentest against a deliberately vulnerable demo site from the HackZero interface (also available over MCP). Pick black box (just the URL) or white box (the source), and optionally have it open a pull request with the fix. Findings are AI-reviewed or reviewed by a certified human tester, and you can supply role credentials (doctor, patient, admin) so it tests authenticated paths, where many real vulnerabilities live.

Jump to 8:00

Why an attacker beats a coding agent

Telling Claude or Codex to make this secure is not the same as an agent that behaves like an attacker: spinning up a browser, hitting endpoints directly with real curl requests, and proving a hole exists rather than assuming the code is fine. The edge is the tools and the harness, which are also exposed over MCP. Internal benchmarks are coming.

Jump to 10:45

Replay: know instantly if a fix held

Normal pentesting is find, fix, then pay for a whole new test to confirm. HackZero replays only the exploit, so you know immediately whether the fix held, and you can dismiss false positives with a reason. Anyone who can verify domain ownership with a DNS record can run it, including consultants on existing contracts.

Jump to 13:45

Why AI-written code is quietly insecure

Models were trained on the whole internet, insecure code included, so they reproduce human mistakes and occasionally leave backdoors. The founders cite a study where models asked to build a simple auth system left it insecure about 45 percent of the time, and with roughly half of new GitHub code now AI-generated, vulnerable code compounds every month. They tease a zero-day found with HackZero in a large public repo, to be disclosed.

Jump to 17:15

Where black box shines: misconfigurations

Platforms like Supabase and Firebase are only secure when configured right. Leave row-level security off, or Firebase storage world-readable, and the black box finds it from the outside, the same pattern behind celebrity file leaks. It is not about the tools, it is knowing how they get misused.

Jump to 20:25

Emergent behavior, and who it is for

They have to check their own agent did not cheat a benchmark; in one run it tried to phish its way through. The close: security is for everyone, it is free right now because they want feedback over revenue, and it is most useful for startups that cannot afford a $50k manual pentest and for regulated industries (HIPAA, SOC 2, PCI, ISO) and fundraising.

Run the same test on your app.

Verify your domain, point HackZero at it, and get the same exploit-validated findings. $299 a month under 10 people, $499 above, with an independent CPA for the SOC 2 attestation.

Full transcript (HackZero segment)

Auto-captions, lightly cleaned for readability. Meaning unchanged.

Show transcript

Ladies and gentlemen, we are back again. We're going to continue this theme of security because there are AI agent swarms that are starting to go out there and they can possibly start attacking you. But how do you check for that and spin up your own stuff to do that? And that can be a little bit complicated. And so today I have the folks from this place called HackZero.ai.

and so for HackZero.ai they're basically allowing you giving you all these tools that you can use including like an MCP tool server that will spin up some agents and they have like a really cool user interface so we're going to learn a little bit more about like security and testing your products from a black box perspective because today you can just give it a url and it can start to go after and start to do that type of testing for you and if you've ever entered this space and try to figure out and do that stuff, especially if you're trying to do some compliance work like SOC or HIPAA and so forth, that could be quite cumbersome because the first thing you'll see on these people's websites are just saying, sign up here and then let us chat with you, and then that could just be a really long process. And I met these guys at the Cursor Compile event, and so I was at the Cursor event, and if you stay tuned towards the very end. There were some folks who were on earlier and we were talking about putting together the special keyboard that the Cursor folks gave to us at the event. It's a Keychron keyboard. It's RGB light up keyboard.

We're going to assemble this today with the custom switches, with a custom keycaps, even a special Cursor custom keycap. So we're going to get that going on today. So without further ado, I don't want to waste any more time because these guys have been holding on through all of my different issues that I've been experiencing. We have the guys here from Hackster. So go ahead and introduce you guys.

Give us a little bit of background of where you guys are from, some of the award-winning stuff. You guys are a really big deal in Latin America. I think I'm trying to bring you guys under the scene and let people know what's up with you guys. So welcome to the show. Well, thanks for having us, really.

It's really a privilege to be on your show. We're really fans of you, so that's pretty cool. I'm Ryan, I'm CEO of Hack Zero. I'm Cuau, founder. Yeah, and a little bit of some background by myself.

I spent the last decade doing competitions in cybersecurity, like CTFs and all sorts of stuff. Recently, I think last year, we got first place in Mexico, second place in Latin America, and it was really fun. and more and more as time went on I started using more and more AI into like you know these competitions and turns out AI is pretty good at hacking you know so yeah that's some intro on myself. Yeah so I've been doing like technology and startups for 14 years. I ran through Y Combinator, Techstars, I did marketplaces, I did niche Software Development, and I also did hardware.

So I really, really love building, and the last two years I've been full heads-down building with agents. So we know what you can code now super fast, but we also know that together with that power, comes a lot of vulnerabilities. With great power comes vulnerabilities. Yeah, with great power. Also, with a great mustache comes a great product.

Our man has debated this. Great, people are loving the stache. So let's talk a little bit more about what you guys are kind of offering here and maybe get a demo as well. Because, you know, like this product that you guys have built called HackZero. Yeah, tell us more about this.

Yeah, so we basically spawned AI agents to hack your company, and then we tell you how we hacked it and how you can fix it. It's got actually quite a few features. It can do PR reviews. I mean, PRs, once it finds vulnerabilities, it can do a PR and it can go straight into your GitHub, so you don't have to spend time into like, hey, okay, I got this report that I paid this dentist company for, and like, what do I do next, right? A lot of the people we talked to were like, hey, okay, we paid this company and now what do we do?

So we kind of sold that. The greatest thing about this is you don't even have to use the platform that we have. We have an NCP. It's a great NCP. So you can just tell it, hey, go find more of one of the radio ladies and then let me know if they're actually fixed.

Should I share a little bit of my screen to show? Sure, yeah. Yeah, let's just share a screen. So you guys haven't even done a formal fundraise. This is just straight out of your guys' garage, in essence?

Yeah, yeah, yeah. We haven't started our fundraising. We already have some angel checks and some interesting investors, but we just got started like literally, how many weeks ago? Like three weeks ago. Like three weeks ago, yeah.

Dude, how many people have signed on to the service or how many people have you got to try it out yet? Well, right now it's free for everyone to try it. So a lot of people that have tried it so far, they really like the product. of course there's some friction that we have to solve but we're still just getting started. We're working with a Mexican bank right now and other people that we met in San Francisco at the Cursor event that just tried it.

So that's awesome and they could just go to hackzero.ai they could start reaching out to you. I mean I feel like if you're in the investment scene right now too don't sleep on these guys because yeah like they got some crazy background and yeah I'll go ahead and see if you can load up your demo and I think I can I should be able to try getting load up on screen. and then I'll switch over to that screen once it gets loaded up on there. So we'll get that going. Yeah, go ahead.

Yeah, you just can see my screen, right? Yeah, yeah. Yeah, so this is what our interface looks like. Again, you can use everything in the MCP, but we're just going to do it here in the platform. So you can just, we made this little webpage called a demo bio site.

It has a lot of vulnerabilities. Just test it out and give it like a quick demo. And you can just click, it's very easy. You just click Run Pentest and you can select white box or black box. for more context for people that don't know.

Blackbox, we just give us the URL and we'll pack it. Whitebox, you give us the source code. And you can actually choose to generate PRs as well once it finds it. So it's pretty cool. We have two options, human reviewed or AI reviewed.

For human reviewed, we actually get somebody that has full certifications or CPS. And they can check manually and perform even a longer pen test themselves if you don't want full AI dimension. but anyways you can continue here select your repository here we select the one you can put credentials so for example if you have like a hospital system you can put like what the credentials for a doctor for a patient for an administrator and you can check that all then should do exactly what they should do because sometimes vulnerabilities don't even come from like unauthenticated users but yeah you can put it there it will ask you if you choose to skip it some instructions if you want to give it like you know something that you wanted to really drill down on and then you can choose when like start now right and then it can just like start depend test and depending on like how big your repository is or not it will take you know anywhere from one hour to six to twelve hours so it can take a while but it's great it will just found like hundreds of agents with different like it will if you give it like the doctor and the patient credentials it was fun like a hundred with doctor credentials and 100 with patient credentials and they will just try to like break things. We're not going to wait six hours here so we can just skip straight into looking at the findings. We do have a question from the audience which is please explain how your solution is different from something like codex security.

I think this is a really good question to ask because I know a lot of people are using the built-in like slash commands and other stuff. Yeah, that's a great question, right? Because you can use still Claude, right? Or you can use still Claude. We did run some internal benchmarks and it turns out like if you give it the right tools and you set up the hardness in the right way, it performs much better than it could with Claude or with Claude.

So it really comes to like the tools we give it that are different than the ones that you currently have on your machine unless you set it up correctly. So yeah, we're going to be releasing some benchmarks soon. Yeah, and also, So, for example, when you give instructions to the AI to make it secure or avoid these kind of vulnerabilities, it's very different from spinning up a browser and trying to click around and pointing directly into the endpoints and having real agents trying to find vulnerabilities directly into the endpoints. So basically, it's like an attacker, not like on the side of the defense. You first attack it, you find the vulnerabilities by really digging into really doing the CURLs or the requests and not only thinking like the agent, hey, this is secure and that's all.

Basically, it proves that it is or it isn't. Yeah, it comes out to the tools. And one last thing. Oh, sorry. Oh, I was thinking about your backgrounds because you have background in security and testing and these additional instructions that you give to your agent harness are also exposed over MCPs.

Like you're saying, if you're using ClawCode, if you're using codecs for the security review, loop in the hack zero part and then you'll get even more comprehensive, actual, like as if you hired a PEM tester, you know, or someone took work. But like they're saying, the agents will then spawn off into all these different domains. domains of testing. Yeah, and the most important thing is if you tell Claude or Codex to keep finding vulnerabilities, it will just keep finding things that aren't really vulnerabilities. Like I'm sure like most of you like have run this experiment of like telling me, hey, find security holes, find security holes, and you just keep finding things that aren't really like worth fixing.

So what we have that, you know, Codex doesn't have is we have like this replay problem. So every single vulnerability that we port, you can actually just like, hey, okay, you found it. Is it a false positive or not? Okay, just replay the vulnerability. Oh, can you make the font a little bit bigger so that we can kind of see some of this stuff?

Yeah, yeah, perfect. Yeah. Yeah. So we have this replay button right here and you can just click it and it will tell you, hey, I'm going to do this, this, this, this. And then it will do it and tell you, hey, you fixed it or you didn't fix it.

Yeah. This is really cool because generally when you do pen testing, you find vulnerabilities, then you fix them and then you do another pen testing so you know that now they are resolved. But with this you only replay the vulnerability and it replace the exploit and you already know if you're still vulnerable or if it is fixed. Oh wow. So if maybe somebody already works in a security industry, can they use your product and test that and like be a consultant or something too?

I feel like someone who has already contracts. They could just keep running this, you know. Yeah, I mean, as long as you can verify the domain, that's one of the steps in running a Pentest, which is just like putting a record on your DNS. As long as you can do that and prove that you own it, yeah, you can just use HackZero. I see.

Okay, yeah, because this is very powerful stuff. I mean, first of all, just being able to replay that once you make the changes in your code, just to make sure that you fixed it is really important because normally when you hire people to do this, they'll charge you again to do the same run again and so you'll get a report back whether it's fixed or not. You're like it's such a slow process right? Yeah. And one last thing another thing if you sometimes things that are vulnerabilities are not really vulnerabilities they're like false positives like yeah I did this thing and it's supposed to be like that.

So what this does is you can just say like hey this is a false positive because of this reason right and next time it runs a Pantast and he finds a similar vulnerability it's going to be like Hey this was a false positive last time so it probably going to be a false positive this time so I not even going to show that And it going to understand your app the more you use it So it your own like chief security officer that alerts your whole code base the more you use it Okay So there a lot of different things to unpack here. One is that when you run an agent run, you start to have things that are very specific, like you say, false positives, tests that you want to repeat again that are very valuable to make sure that you fixed them. All of these things you'll want to somehow capture. And your platform will basically keep that for them. It really exposes it over MCP.

Agents can start talking to it. They can do these round trips. And it really kind of takes advantage of that, but also approaches it from an actual security researcher background type of thing where it's doing various vulnerability tests. One is literally just going to your website and testing as if you're an external user. but the other part of it is to go deeper and say like give access to the code base and then it itself will then start to like go through things and make sure that the security bugs it is finding is actually of higher signal and not just like a hallucination, not necessarily a hallucination but maybe something in this training set.

Can you talk a little bit more about these models, right? Because I feel like low-key, they may be leaving vulnerabilities maybe because of how they've been picking up all the, like, open source code that it's been trained on sometimes has vulnerabilities left in because three-letter agents leave them in there on purpose. Like, and so if they see enough of those examples, they may accidentally make a backdoor into an app whether you prompted it or not, right? So tell me more about how your software can also kind of look for these types of things. Yeah, so like you said, like, you know, we didn't get ChatGPT or all these great models just with no data.

You know, we had to train it on the whole internet and, you know, repositories and code, right? And humans make a lot of mistakes. That's what makes us human, right? And these mistakes are, and not necessarily mistakes, but sometimes, you know, we saw a study, I don't know if you remember from who, They ran like 80, they told AI like different models to make like an authentication system, like a simple login or like different things, like 80 of them. And it made, it didn't care about security 45% of the time.

It just like completely. Yeah, that's one vulnerability at least on 45% of the generated code. And that's generally what you will ask with AI. AI, and as you know, right now on GitHub, I think almost half of the code has some AI generation. So basically, every single month, we're generating more code with AI.

So this is growing exponentially. So we need fast solutions because we're getting a lot of vulnerabilities every single month. We get more and more. We also, we discover one, like, zero day. We will announce it later when we can in a big public repo with Hack Zero.

We cannot say because, but we will announce it soon, as soon as we are able to do that. But it's interesting. It's interesting. We're shipping a lot of code and also a lot of vulnerabilities. Okay, so someone is like in a highly regulated industry like healthcare, law, things where you're charging and you do require some more security scrutiny.

You know, people are recognizing, well, like this isn't a cheap product, but if you ever do type of testing like this for security, it's 15 to 20K typically for, I mean, these engineers get paid a lot of money because it involves very critical systems. If you lose data for all these customers, you're like in these million plus dollar lawsuits, you know, and charging 15 to 20k for like one run is like kind of nothing to them. And how should us as builders who may not be at that level yet also think about security? It sounds like there's, you know, many people may not be aware that it could be generating some things. I like to use some of the stuff that's off the shelf.

Like I use, you know, clerk for my authentication. You know, I'm not trying to write my own auth from scratch. I use established providers like Convex and they already have these working relationships and they have engineering teams who are already kind of solving these things that's kind of some of my best practices, I don't try to roll my own auth from scratch because I know there's a lot of edge cases and I don't even try to get to them but is that the right way to do it especially if you're just starting out prototyping and building just so you get downstream a lot of things, I mean what do you guys think? Yeah, so you don't really want to reinvent security from scratch, right? if you want to use these open source things that have been audited for a while.

But regardless, this is where Blackbox really shines because, yeah, these apps or these different platforms or vendors are giving you these platforms to do this, but they sometimes require you to configure something. And if you misconfigured that thing or you didn't know that you could even configure that thing, the Blackbox will just find it. So if you disable like row-level security or something like that on Supabase, you know, Supabase is secure. just didn't enable it or, you know, and the black box just found it. So it's not really using the tools but knowing how to use them and this is where our product really shines.

okay, that's a really good point. So like even just for doing one run, you can see, oh, I didn't, like Firebase, it left all of the user files so that everyone can see them. And it's kind of a repeatable pattern that an agent can figure out to grab everyone's files by running a script. Yeah, well, you know, maybe I was too lazy two months ago setting up my Firebase rules, and I just left it like everyone can read and everyone can write, and I just forgot about it. I just started getting more and more users, and then, you know, like I'm Marcus Brownlee, and, you know, my wallpapers get leaked, and so, you know.

I see. So it'll find stuff like that that maybe the AI agents will or will not find if you prompt them, but they kind of have to think in this very special way. This is a really good question that someone also was talking about here Can you give a rough estimate of how much better HackZero is compared to other vulnerability scanners like Nessus or OpenVAS? Yeah, so I guess better is an interesting word But everybody covers different things I think some of these vulnerability scanners, they have a lot of templates for things but sometimes they can get like very limited or like if you have like a firewall like We can tell about the benchmark. We've been running the Expo benchmark.

Yeah, I think we should probably hold off on that a little bit. Yeah, but yeah, sometimes you need you have like a firewall or you have things that are blocking you to do that. So you're going to send like a thousand requests and it's going to take a while, right? And you have to like download everything and have the right templates. AI is really smart and knows really how to optimize.

I think hacking is really just search at the essence of it. And LMs are really good at searching and knowing what failed and what didn't fail. So if you start running, I don't know, if you run the template for PHP and your thing is like Django, you're just going to waste a lot of time. And we'll figure out, oh, you're using this stack from the beginning. You're using React, I'm going to search for.

Yeah. Thanks. And talking about the benchmark, we still are making sure that the numbers that we got are real because for us, it's like it looks a little bit unreal. So we are testing that every single pass was really a pass because sometimes our AI escaped and it downloaded the code and spin up a Docker and then did everything like cheating, you know? So we need to make sure that every single pass we have, our AI didn't cheat it.

Because we also had an instance where it was emailing us or trying to do phishing also. That's crazy. Yeah, you guys are going to get peer pressured or something like that. That's amazing how stuff is kind of kicking off. Wow, wow.

Any other things that you feel like, I mean, because they're super early and people can kind of try this out. Who are the people that you're really trying to reach the most that would be really insightful for you guys to reach to at this stage? Yeah, so security is really a thing for everybody, right? Right now, if you are like a startup and really can't afford to pay like 50k for a pen test and you really care about your security and you don't want to get a bad reputation because, you know, somebody hacked your thing, you should probably try it out and it's free. We really are just looking for feedback on our product, trying to reduce friction and trying to like if everything if anything didn't work or you didn't like something or we couldn't prop something that is insanely more valuable to us than money right now.

So, yeah, just try it out and let us know what you think. So it's free. They just put their domain name and say get started and it'll just get free test. Yeah, you have to do a little like more steps like verify your DNS settings or like whatever. Also, we recommend a lot for regulated industries because generally it's very expensive to do manual pen tests.

So for HIPAA compliance, SOC2, PCI, or ISO, it's very helpful. Also, if you are fundraising, you know you can save some money by doing AI penetration testing instead of manual pen testing. That's amazing. That's pretty cool, especially if a product that's going out and it's picking up a lot of progress and people are wanting to do that. I think that's definitely a cool way to go.

And I'm really happy to have just bumped into you guys at the Cursor Compile event. It was a really, really fun event to kind of hang out and kind of see what they're cooking with. I mean, they got the GitHub Origin thing going. They got the competitor, they have the new model and a lot of really cool stuff going on. So I think, yeah, on that note, I think what I should probably start doing is put on my little compile hat and as promised for people, yeah, right, yeah.

Did you guys get, you guys, oh, you did get one, you got a cursor one, right, yeah. Let's go. Oh, the other way, yeah. Awesome. Oh, you got the pins too, yeah.

Yeah, that's just cool. Yeah, I also got my little codex keys, you know, my little keys. Oh, yeah, those are awesome. This is how, this is how I use, I code, I just go like this. Yes, yes, yes.

Tap. Tap. Yes, yes. Continue no permissions. Continue no permissions.

YOLO! Insert all of the bugs. Yeah, that's the violence. You guys got ripped out. Wait, did you guys get a keyboard too?

Yes. Yeah. What's your config? Did you put it together or what colors and stuff? No, not yet.

We're actually going to join a group where everybody's going to do the keyboards at the same time and do networking. It's going to be in a couple of hours, yeah. If you want to come by. Wait, where? Yeah, I can see the link.

It's going to be Korea versus Mexico, soccer thing, whatever, and everybody's going to build their heroes. Wait, it's Mexico versus who? Korea. Oh, Korea. Oh, wow.

Heck yes, dude. Wait, is this in San Francisco? Yeah. Okay, I'm right now in the South Bay. Yeah, what's up?

Yeah. That's funny. Maybe I should do an IRL stream from there. That'd be crazy. Police is coming.

All we're going to hear all day is going to be like, no! So how is the Korean team? I have no clue. What do you think? No idea, man.

What do you think? I never watch soccer. Whenever it is like the World Cup, that's what I drew from Mexico. Exactly. Yeah, I was walking by the streets of Mountain View and I saw all the Colombian people all dressed up.

They had the horns and everything When they made the first goal it was so loud on the streets It was crazy, people were watching all over in the bars and stuff It was funny, because I was like, oh that's right, the games are being played right now I was like, who's playing? So that's cool I just going to continue this stream And just going to do mine to like for those who are watching I know like we had