
What We Learned From Running an AI Hackathon
Sofia Kuttner Lindelow, Forward Deployed AI Engineer, and Sophie Benghozi, LEARNING AND CONTENT SPECIALIST
The goal of AI fluency on a mixed team is not more people who can discuss models and demos. It is people who can build systems for work they already own, since domain experts usually know what those tools should do better than the most technical engineer in the room. Conceptual training can explain what a tool is, but it rarely forces the decisions that make a tool useful in ordinary work, and it rarely produces the evidence that changes what people believe they can build.
We ran an internal half-day hackathon to test protected time to build against a real company or client problem. AI fluency grew when teams had that time, a short proposal the founder had approved, and enough friction left that people had to figure out sources, limits, and quality bars themselves. The rest of this piece covers the conditions we set, why a real problem beat a generic tool, what four teams shipped, and the sequence we would use next time.
Protected time produces builders, not only users
We announced teams about a week ahead so groups of two to three could pick a concept that was feasible but still a challenge. Each team wrote a short proposal, and the founder approved those proposals before build day. A helper agent opened by interviewing the team about the problem itself, which kept the session from collapsing into a tool tour. We blocked a half day for live building, with the option to keep going through the week, and told people to start from a real internal or client pain rather than the most obvious automation of their current role. Our data warehouse and company brain served as the foundational data and context for all builds.
About a week later, four working tools were presented: a landing-page generator, an onboarding Slack agent, an email campaign workflow, and an analytics build for reminder, incentive, and win-back windows in order history. Importantly, people who do not typically build production systems every day underestimated what they could ship under protected time against a problem they already faced.
A real problem forces decisions a generic tool never requires
Each team started from a real pain point: new-hire role questions in slack, the monthly repetitive process for full bilingual email campaign calendar to scheduled send, the time cost of assembling landing pages from a brief, and a year of unused order history. Once the problem belonged to the team, someone had to name the input, the source of truth, what counted as a failure, and who decided when the output was good enough. A generic training risks walking through features without practicing implementing those guardrails.
The onboarding agent made that distinction concrete. It answered new-hire questions by pulling from the open web instead of internal documents, and the answers could sound fluent while citing the wrong source. Retrieval stopped being a concept and became a decision the team had to understand: which corpus is allowed, how a wrong source shows up, and what to do when the answer is confident and still wrong.
Keep enough friction that people have to learn
The teaching sequence we kept seeing was try, hit a problem, need knowledge, learn, test, and adjust. Most training assumes the inverse path of learn until you feel ready, then try, and that path is slower at building the capability we want. Credits ran out on the landing-page build, meaning the account's usage allowance was gone. Slack failed to connect for the onboarding agent, wrong starting instructions from another agent sent a run down the wrong path, and a token limit stopped another run before it finished. Each blocker required a concrete fix that later becomes ordinary judgment.
The credit failure also taught a placement lesson. The landing-page work lived as a skill, meaning reusable instructions kept in the shared workspace, and the build survived a model switch with those instructions still in place. Two days later a second version existed as a reusable skill. The organizer job is to stay reachable so a stuck team gets an answer quickly, and to leave enough friction that people practice figuring things out.
Four builds as evidence, not a product tour
The landing-page generator takes a short brief and an optional image and produces store theme sections with editable images. It ran out of credits, survived the model switch because the skill lived in the shared workspace, and reached a reusable second version two days later. Automatic checks and bilingual pages are still open additions.
The onboarding Slack agent was meant to answer new-hire questions from internal context. It answered from the open web instead. Slack failed to connect, wrong starting instructions derailed a run, and a token limit stopped another, yet a scope test still worked: asked for work outside its job, it refused and escalated.
The full-stack email automation takes a one-line brief to eight bilingual emails for Klaviyo, with copy, images, and links. This meant combining a fleet of skills for copywriting, QA, image generation, and integration to our OS for seamless use. Initial runs were great prototypes, but building it for sustainable production took countless trials.
The analytics team used about a year of one client's order history to find reminder, incentive, and win-back windows. At the October 1 show and tell the analysis was running on two laptops. Next comes hosting for the whole team and a direct connection to Klaviyo.
A working demo and a tool people rely on every week are different standards. The four builds proved the ability to build a functional MVP is in the hands of the user, and has given every team member the actual skills to continue developing their own solutions.
How to run the next session
1. Block the time in advance, and explain the gain for a non-technical role by saying what that person will be able to do in their own work after the session.
2. Start from a problem the team already has. A helper agent can begin by interviewing the team about that problem, which is what happened here, so the session starts from the work itself.
3. Review a short proposal before the build. The founder approved those proposals for this hackathon.
4. Mix experience on the team, and name who to ask when someone gets stuck, whether that is a person or the helper agent.
5. Let the building continue after the scheduled block. People had the option to keep going until Friday, and I would keep that option on the calendar.
6. Finish with a show and tell of what works and what does not. Ours was on October 1 and covered the four tools, including the connection that failed and the runs that stopped.
7. Test every tool before calling it finished, using a case the team already recognizes, the way the onboarding agent was given work outside scope.
8. Plan the next session before the group goes back to ordinary work, and give it a defined outcome. If two teams try two approaches to that same goal, the comparison will be easier than it was with this open brief.
Readers will need an equivalent of the shared AI workspace we already had, with brand material and warehouse access in one place.
Repeated chances beat one-off training
AI tools will keep changing. What compounds is a team that has practiced solving real problems under protected time, kept enough friction to force real decisions, and shared what worked and what broke. The near-term test is which builds still get used, and what those same people bring to the next half day.


