Internal AI Adoption Hackathon: What We Learned

Internal AI Adoption Hackathon: What We Learned

By Antonios Papadopoulos

Overview

Anvil Secure's Technical Director, Antonios Papadopoulos, breaks down how the company ran a remote, company-wide AI Adoption Hackathon: two cohorts, dedicated build days, tracks drawn from real work, and strict human-in-the-loop guardrails baked into the scoring rubric itself. He shares what came out of it, and a practical checklist for any security team considering the same exercise.

Most security people are skeptical by nature, which makes us a tough crowd for the ongoing AI hype. Rather than keep debating it, we ran a company-wide AI Adoption Hackathon instead. Two remote cohorts, engineers in pairs, one protected build day each, strict human-in-the-loop guardrails and, at the end, a triage of real projects to adopt, incubate, research, or shelve.

Why a hackathon and why now?

Like everyone in security and beyond, we have watched AI go from a novelty to something people reach for daily and for almost every task. For us, the question was never really whether to adopt it; it was always about the how, and how we do it without losing our identity or the standards and the quality our clients expect from us. And, frankly speaking, there is only so much you can learn from theory alone. Without getting your hands dirty, it is hard to understand the possibilities and use cases; especially when jumping constantly from one engagement to the next.

The idea is not novel. Earlier this year, a blog post from one of our competitors on how they made themselves AI-native made the rounds internally and became a recurring topic in the conversations that followed Hammercon, our annual internal conference. It became clear that equipping our engineers with subscriptions to LLM providers and giving them access to frontier models alone was only step zero. The real adoption would come by getting our hands dirty and using the tools to build, fail, and repeat. Failing fast and often is one of the better ways for learning effectively, and a hackathon was our way of facilitating this.

So, we kept the premise simple: give people a real and uninterrupted day, real problems from their own work, and clear rules going in, then see what they come back with.

How we ran it

As a remote-first company with people residing all around the world, running a hackathon in-person would have been a difficult undertaking. Instead, we decided to run it remotely splitting it into two cohorts: one for our Europe-based engineers and one for the ones based in the U.S. We worked in pairs, each team got a dedicated build day, and everyone had the same tooling and the same rules going in.

The part that took the most thought up front was the tracks. We didn't want a free-for-all where the vast majority built the same red teaming tool, but we also didn't want a list so rigid it boxed people out of their own ideas. Instead, we started from the work itself and asked a simple question: where do we actually spend most of our time, and where could AI help?

That gave us a handful of focus areas: our core testing work, the internal platforms we build and maintain, the project-management and delivery side that keeps engagements running, and the particular client ecosystems some of us work in every day. Each became a track. We also created an open "wildcard" track on purpose, so a good idea that didn't fit any of those boxes still had somewhere to go. Looking at the end results, most of our engineers decided to work on the predefined AI-assisted pentesting track and develop tooling that could be leveraged for future engagements. Given our engineers' predilection, this was not really a big surprise.

To keep things comparable, every team shipped the same way: working code committed by the end of the build day, plus a short write-up of what they made and, just as important, how they used AI to make it.

Guardrails, on purpose: keeping humans in the loop

Now here's one of the parts we care most about. A security company adopting AI has to hold itself to the standard it sells; it is the same scrutiny we bring when we test AI systems for our clients, and it would be odd to hold our own adoption to a lower bar. Our COO, Dana Hehl, wrote earlier this year about how Anvil approaches AI: AI as a tool that supports expert judgement, never replaces it. The hackathon was our chance to put that principle to the test.

At the same time, it's worth being clear about one thing. A lot of the current conversation treats the phrase "AI is just a tool" as a cautious first step. Essentially, something to grow out of on the way to going fully "AI-native." We don't see it that way. The risks are well documented (prompt injection, data leakage and the rest of the OWASP Top 10 for LLM Applications) and they do not fade as a team gains experience. For a security firm, the guardrails are not training wheels we take off once we are comfortable. They are even more important than the end result.

So the rules were explicit from day one: humans stay in the loop, no unreviewed AI output goes anywhere near production or client work, and nothing touches confidential data without explicit approval. None of that was fine print; it was written into the rules and scoring rubric, underlining the importance of being intentional with the usage of AI.

If anything, the constraint sharpened the work rather than slowing it down. The projects that stood out weren't the ones that handed the most off to a model; they were the ones where AI made an already-skilled person noticeably faster, while that person stayed firmly the one making the calls.

What we took away

We were deliberate about judging for more than a polished demo. Every project was sorted into what came next: something to adopt, something to keep incubating, something worth treating as research, or something to quietly shelve for now. That distinction mattered to us; frankly, a hackathon that ends the moment the last presentation does is just theatre. The point was never a fun day away from client work; it was to come out the other side with a pipeline of things we might genuinely use. Equally important was the opportunity to learn how to use the tools available more efficiently, identifying what works and what doesn't.

Outside what the end results and projects were, what really stood out to me was watching people who spend their days taking things apart get genuinely excited about putting something together and, just as importantly, doing it carefully. Curiosity with a bit of discipline behind it; that is exactly the combination we want shaping how Anvil approaches AI.

Looking ahead, we are already thinking about where the strongest ideas go from here. The ones we marked to adopt will receive additional support and dedicated building days to convert from prototypes into actual tools and applications that can transform the way we work. We expect that a subset of them will naturally evolve into presentations and possible publications. Whether that ends up on our blog, at Hammercon, or at other conferences is something we will figure out along the way.

We walked in as a room full of skeptics and walked out with a list of things we actually want to keep building. For a company of professional skeptics, that is not a bad place to land.

How to run your own AI adoption hackathon

If you want to try something similar, here is what actually mattered. It was less about the mechanics, and more about the things that would have broken it if we had got them wrong:

  • Protect the build time, for real. A dedicated, uninterrupted day per team is the whole game. That means ring-fencing it with project managers well in advance, not just blocking a slot in a calendar. Make sure no planned meetings and interruptions are present for that day, to the extent that it's possible.
  • Protect the judges' time too. Judging is real work, and it is the part most likely to get squeezed. Give judges a proper, separate window to read the submissions, run the code, and actually think.
  • Draw the tracks from your real work. Start from where your people already spend their time and where AI could genuinely help, then add one open "wildcard" track so the ideas that don't fit still have somewhere to go.
  • Put the guardrails in the rules, not the footnotes. Humans in the loop, nothing unreviewed near production or client work, no confidential data without explicit approval and make them part of the score, so responsible use is rewarded rather than just assumed.
  • Ask for something real, then follow through. Working code plus a short write-up of how AI was actually used, judged on one shared rubric and give the projects you choose to adopt real time and ownership afterwards.
  • Plan thoroughly, then expect the plan to bend. However complete the plan looks on paper, something will drift. Dates move, people get pulled onto client calls, some people go on vacation, among other things. Do your best to protect the few things that genuinely matter (the build day above being the obvious one) and hold the rest loosely; the real skill is having planned well enough to know which is which.

About the Author

Antonios Papadopoulos headshotAntonios Papadopoulos is a Technical Director at Anvil Secure, where he focuses on web and mobile security and enjoys streamlining processes and mentoring teams. He organised and ran the company’s internal AI Adoption Hackathon.

Tools

aqlmap - A tool to extract information from ArangoDB through AQL injection. See the introductory blogpost.


awstracer - An Anvil CLI utility that will allow you to trace and replay AWS commands.


awssig - Anvil Secure's Burp extension for signing AWS requests with SigV4.


ByteBanter - A Burp Suite extension that leverages LLMs to generate context-aware payloads for Burp Intruder. See the introductory blogpost.


dawgmon - Dawg the hallway monitor: monitor operating system changes and analyze introduced attack surface when installing software. See the introductory blogpost.


GhidraGarminApp - A Ghidra processor and loader for Garmin watch applications. See the introductory blogpost.


HANAlyzer - A tool that automates SAP HANA security checks and outputs clear HTML reports. See the introductory blogpost.


IPAAutoDec - A tool that decrypts IPA files end-to-end via SSH. See the introductory blogpost.


nanopb-decompiler - Our nanopb-decompiler is an IDA python script that can recreate .proto files from binaries compiled with 0.3.x, and 0.4.x versions of nanopb. See the introductory blogpost.


OffTempo - A Burp Suite extension for statistical timing side-channel analysis. See the introductory blogpost.


PQCscan - A scanner that can determine whether SSH and TLS servers support PQC algorithms. See the introductory blogpost.


SAPCARve - A utility Python script for manipulating SAP's SAR archive files. See the introductory blogpost.


ulexecve - A tool to execute ELF binaries on Linux directly from userland. See the introductory blogpost.


usb-racer - A tool for pentesting TOCTOU issues with USB storage devices.

Recent Posts