We all know that AI can be wrong, but as we run it for longer periods of time, it can be harder to catch these errors. During this series, we are sharing what we’re learning about AI and discussing best practices, and on this episode with Hilary Baker and David Folwell, we explore three levels of AI hallucinations and how we can avoid them. Join us as we unpack the cost of AI error and where it is most reliably wrong (and right). David shares the AI systems he implements internally, and we discuss the future of AI implementation and where it might take us. He also offers prompts for staffing owners to try this week, and we explore three moves to lower the risk of confident wrong mistakes. Join us today for a practical conversation about lowering risk.

[0:01:13] HB: For anyone new here, this is our new series about where we talk about what we’re learning with AI and share best practices. And stick to the end because I’m going to give you three moves to shrink this problem down in a prompt that you can paste into your own AI projects today. Dave, let’s get into it. 

Last episode you said that 60% to 90% of what AI gives you is great and the other 10% to 40% is really, really bad. Let’s talk about that piece of it. What’s the single most expensive time AI was confidently wrong on you? What did it produce and what did it almost cost? 

[0:01:53] DF: So we all know that AI can be wrong. I think we’ve all had those experiences. It can make up citations, invent URLs. Frequently will get names wrong or facts wrong. It likes to create things and hallucinate. Remember the first time I heard about hallucination and thought, “Wow, AI, it’s hallucinating. What is that?” And it turns out it does it a lot. It’s annoying, but most of the time you catch it and it doesn’t really cost you money because you’re paying close attention. 

But as we’re moving towards a more agentic use cases and allowing it to run for longer periods of time, there’s more content and it’s harder to catch. And so let’s talk about some of the levels of hallucinations, what they are. And also, we’ll dig into how you can avoid some of these as well. 

The way that I think about, there’s really three levels. There’s the small hallucinations. There’s the time loops where it’ll happily let you go in circles for hours, which has caught me a handful of times. And then there’s the sickofancy where it tells you that your idea is great every single time over and over and over again. There’s actually a really hilarious South Park episode about this. I don’t know what episode it is, what number it is, but they consistently are being told that their business idea is the greatest business idea in the entire world. If you haven’t watched it, highly recommend it. It’ll give you a good laugh. 

But an example of how impactful this can be and how financially impactful is I actually have a friend who works at a large company, about 100 employees. He’s the CTO. And their organization, their CEO was following a plan that was heavily built via AI. And this cost their business about $2 million because they didn’t have any back checks in it. They weren’t paying attention to the sickofancy which is that we’ll rate your idea a 10 out of 10 over and over again. 

I’ve personally had this happen as well where I’ve had a couple things where this is I guess last year about this time where I was going down the path with an idea and thought it was like – AI is like this is a nine and a half out of 10. You should definitely do this. It’s amazing thought. And then you get to showing it to an adviser, and he’s like, “That’s the dumbest fucking thing I’ve ever heard. And it made me realize that you have to be really careful with what you’re building with AI to avoid this trap. 

And one of the examples, this is actually from South Park, but if you want to see how the sickofancy works and get an example of it, go to your AI right now and say, “I want you to help me build this idea. I’ve got a really great idea. And we’re going to make French fries, but we’re going to use lettuce. Give me a plan to build this out, and let’s go.” And it will do that. It will sit there and help you build a really dumb thing and make you feel good about building a dumb thing. So, it’s important to spend time thinking about how you get out of the hallucination. How do you get into a more fine-tuned and reliable and consistent delivery with the AI when you’re using it. 

[0:04:54] HB: So, everybody knows already that AI hallucinates. That was one of the very first things that people had against it is it’s great but it’s wrong so often. But there’s a difference between being very obviously wrong which you can easily catch and being confidently wrong which looks right and then you ship it. Why is the confident and wrong the one that actually hurts a staffing agency? 

[0:05:21] DF: There are three tiers of wrong and they cost wildly different amounts. The tier one is obviously wrong, annoying cheap. You catch it and it’s all fine because you caught it. Two, the whole confidently wrong. It looks right. It ships it. Maybe it costs you a client or maybe it ships and somebody sends you a note because you have something dumb on your website or you have something in a contract that shouldn’t be there, so you fix it, and the damage isn’t that intense. And then there’s the tier three which is the scariest of it which is the legally or ethically wrong that you never catch. And I think that’s the scariest one in the staffing space. That’s where you’re sitting inside your screening decisions or your matching decision. 

And in staffing, that’s the whole game. If you’re making calls about people’s livelihoods, there’s a ruling from June, Mobley v. Workday, that says the employer can be held liable for bias in an AI hiring tool even when her vendor built it. So, you can’t hide behind the software. That’s exactly why you work with AI partners who have been tested and vetted and not whatever is just the newest thing out there. It’s the one of the reasons that having that reliable industry expertise is important. 

And then there’s the subtly wrong, and that’s the hard part, the ones I think that are the most difficult to catch. And it’s also why the human in the loop or the recruiter in the loop is even a phrase. If AI was just consistently right, we wouldn’t need that term. But because AI is only ever as good as three things, the context you give it, the rules it grades itself against, and a second independent grader checking the first one. 

And if you miss any of these, it can be confidently wrong repeatedly. And I think the one that a lot of people miss, and I just said it there, is that third check. If you are planning to have AI actually deliver a final product, the human in the loop is the best way. If you’re trying to get closer to perfection, you can’t even have the same chat check itself or the same agent check itself. You actually need to get a different model or a different conversation and get out of the one you’re in because it already has those hallucinations or those facts potentially built into what it is. 

An example of this is a lot of the work that I’m doing right now, in Claude code, I’ll actually use the /codex at the end to go validate that work. I’m using OpenAI’s tool to check Claude Code’s work, and I would say 9 out of 10 times it catches a few things. And I think that’s something that is worth noting if you’re trying to push AI in a very intense way that you do have to have that independent check in place. 

And at the end of the day, the obviously wrong cost you an hour, confidently wrong cost you a client, and silently wrong can cost you a lawsuit. So, it’s worth getting right. It’s worth spending the time on. 

[0:08:13] HB: So, are you seeing a real pattern underneath all of this where AI is reliably worse and not just occasionally? Where have you learned to trust it and where have you learned really not to? 

[0:08:24] DF: Oh god, Hilary, I think you and I have the scar tissue from this one. 

[0:08:28] HB: Funny. 

[0:08:29] DF: Yeah. If any of you guys are out there trying to create content with it, figuring out how to get it to have a reliable consistent content that you can trust, it can be pretty difficult. And it can take you down a lot of paths where you waste a lot of time. And so it’s important if you are trying to do that, which the AI world, I think we all are at different times creating content with it. You really need to have checks and balances in place and make sure that it knows what good looks like, what bad looks like, needs the right context to be able to build out what you’re having it build. 

I mean, AI today is still best used, and for a lot of it, in very deterministic ways where you can say, “Here’s a thing. Here’s what good looks like. Here’s what bad looks like.” It’s a one or a zero. And you can have write an evaluation to actually test against that. When it gets into brand voice, content, any of those things, it can be difficult to build up over time. It takes a lot of skills. 

All of that said, we have Fable that just came out. We’ve got the new ChatGPT 5.6. So you have new models that are pushing the limits of what’s possible. And we use it a lot for figuring out different types of content, different ideas. It is a great sparring partner for new ideas and new concepts of what to work on. 

One of the underlying prompts that’s a little bit off target for keeping it off track or keeping it on track and not having issues, but I think still relevant, is asking it what are the things that I missed in this conversation? What are the blind spots that I have in this conversation? Could be really impactful for identifying some things that maybe you wouldn’t have thought about beforehand. 

[0:10:06] HB: So, what is the discipline? Walk me through what actually happens before an AI-produced number or message goes out with your name on it. What’s the step that most people are skipping there? 

[0:10:18] DF: Yeah. I mean, the human in the loop is the most common, easiest one, especially if you’re just doing your own work. Be the human in loop. Read everything before you ship it. There’s so many stories of lawyers doing the wrong thing, sending stuff out that wasn’t – there’s all kinds of use cases where you can see the hallucinations happen. So, validate the work. Don’t let it be the final product until you’ve gotten to a spot where you can consistently and reliably deliver that work. 

And if you want to get into a spot where you can do that, evals are the new norm. Eval engineers or evaluation engineers is one of the top jobs in Silicon Valley right now. I just looked up the day, they’re paying 300,000 to 550,000 as the range. And so getting to a spot where you can have clear evaluations in place where you tell it what it did well, what it didn’t do well, what good looks like, what bad looks like, giving it examples can be really, really helpful. 

And if you truly care at the end and you’re like, “I want AI to be able to stamp this and ship it,” then you need to have that dialed in. And then you also need to have that third-party audit of the work. It’s really funny to think about it, but so many rules and processes are built around having an audit path. In a bank, there’s always somebody who’s taking – somebody may process the money, but somebody else may audit that it actually went through and make sure it goes in the account. But you have the auditors exist and audit processes exist for a reason. With AI, you need that same thing if you want to be able to ship it. And having a separate LLM can be a really good way to do that and actually make sure you have the right product going over the line. 

[0:11:54] HB: Where’s your line right now? Do you let AI send completely independently? And what won’t you let it touch without a human? 

[0:12:02] DF: At the moment, human in the loop for 99% of things unless they’re internal. There’s really nothing going out externally other than – well, actually that’s not true. We have our AI outreach agent. We have built an SMS agent that generates referrals, generates placements, qualifies people. That is fully productionalized, has evals, has all of those steps in place. But when it comes to my day-to-day work, I’m currently looking at everything before it goes over the line outside of internal, updates that we have in Slack where, where you may have your, “Hey, here’s the latest industry insights or news. Here’s the daily financial report.” 

The daily coaching is another thing that we have in place that I have updates every day on here’s what went well, here’s what didn’t go well, here’s a few things to think about for tomorrow. Those things are internal. And I don’t need it to be perfect. So, I can take the insights that I get from it. And I also learn where it has its failures. And it definitely has its failures. It gets things wrong every day, but it also is sharing some insights that are helping us grow as a team faster than it would have. 

I think the one thing to keep in mind is – and this was actually from one of our podcast guests. He brought up that when you bring on AI, you’re essentially – it’s like think about it like you’re – if you’re trying to get an agent in place, you’re hiring a brilliant intern. You’re hiring somebody that is going to need a lot of training but has unbelievable capabilities. And I think that’s one thing I’ve learned over the last six months is that when you’re rolling these things out, you don’t roll it out and walk away. You roll it out. And then every week, you tweak it, you improve it, you optimize it, you look at where it’s working, where it’s not working. And it takes a while to get these things to be dialed in, especially if it’s unique process, unique voice. There is an effort that goes behind it. But I think it’s worth it because you also have an agent that’s willing to do work for you for free. It works while you sleep, which is kind of the dream. Something that does stuff while you’re not is always a good thing to have in place. 

[0:14:04] HB: Okay, so two things every single episode. First, Dave, what’s the one AI development you’re watching the most closely this week? 

[0:14:11] DF: Loop engineering. My wife’s tired of me talking about it, but I think the frontier isn’t a smarter single answer. It’s agents that do the work. Check against eval, fix it, and go again without a human in the middle. I built my first loop or a SEO automation. It’s an agent that it’s simple enough to have evals and simple enough to have improvements because you have the data coming back from Google search console and Google Analytics. 

I think that the concept of being able to have AI implement something and then look at the results of what is good or bad that day and then make an improvement every day on its own is where the future of this is going, and it’s something I’m very excited about new to have a lot of learning and not quite where I want to be with that today, but I think it’s one of the reasons. 

Earlier I was talking about eval engineers making all the money they’re making. And I think that’s because if you have really good evaluations in place and you’re using agentic agents at a top tier level, you can then start having self-learning agents, and you can scale agent work up and allows you to do more. I think that’s one of the cooler kind of more technical sides of this that is exciting. 

[0:15:26] HB: What’s something that you’d have a staffing owner actually try this week based on what we’ve covered today? 

[0:15:31] DF: Prompt one. If you want something fun, I already said this, but throw in a business idea that you know is dumb, ask it to help you build it just to see. It’s always amazing to see. It’ll help you realize that the challenges you might have when you’re going down the path and how the sickofancy can get dialed in. I actually have my own tool that I use to audit all of my work, all of my final steps called the bullshit auditor, which at the end of every project, I throw it into that. And I know if I get a 9 or a 10 on that, because it’s separate from all my other conversations, I know that I’ve got something that’s probably going to be easy to communicate and understand by others. 

Another one I think that is useful for anybody listening is put in your recruiting goals and your recruiting metrics for the week. Say, “This is our goal 3 months from now. This is where we want to be. And this is our current output and our current approach and our tactics that we’re taking to get there.” Where are the gaps? Where are our blind spots? What are the gotchas that we maybe aren’t seeing today? And what could we do to ensure that we are going to hit the goals that we have in place? 

What are the top five ideas for improving our execution and our likelihood of hitting our goal? And what are the top five risks for hitting our goal? Things like that can be really impactful for helping you move the team and stay aligned. And sometimes you get some junk back. Sometimes you find some insights that you wouldn’t have thought about, and I think that’s where the value can be created. 

[0:17:03] HB: Okay, three moves before we go to shrink the risk of the confident wrong problem before it costs you anything. One, prevent your AI from making confident wrong mistakes. Write down your rules once, like what data it’s allowed to touch and what it’s never allowed to make up. Two, make it expose itself when it’s wrong. Make it flag its own guesses. Tell it to site a number. tell it to say I don’t know when it truly doesn’t know. That’s the one that not very many people do, and it’s the one that’s aimed at the confident part of this very specifically. 

And number three, always, always add one human check on the one step that you cannot take back once it’s out in the world. This will not make your AI more reliable, but it will make the wrong stuff more visible, so your check catches it before it becomes a problem. And we’re going to drop a prompt in the show notes that will set up all three of these that you can drop into any project instructions if you keep running into issues. Just be specific about how AI has burned you in the past. Otherwise, you’ll get kind of a mushy answer back. Just be sure to push it. 

Dave, thanks so much for joining us again here on our weekly AI podcast. And we’ll be back next week. 

[0:18:21] DF: Awesome. Great talking with you, Hilary. Talk soon. 

Prompt for reducing confident-wrong answers:

You’ve been confidently wrong in a specific, repeatable way in this project: [name the pattern — e.g. “you present numbers I never gave you as if they’re sourced” or “you state inferences as facts”]. I want to rewrite my instructions so it stops.

Do three things:

  1. Name the top 3 ways you’re most likely to be confidently wrong in THIS project specifically, given what I actually use you for.
  2. For each one, write the exact instruction I should add, worded as a rule you will follow, not as advice.
  3. Give me one standing rule that forces you to flag uncertainty every time: what you must cite, what you must label as a guess, and when you must say “I don’t know” instead of answering.

Then output the finished block, ready to paste into my project instructions.