Showing posts with label artificial intelligence. Show all posts
Showing posts with label artificial intelligence. Show all posts

Thursday, July 2, 2026

Auditing a system that uses AI

We’ve talked before about some of the Quality issues related to the use of artificial intelligence (AI) technology, but mostly in an incidental kind of way.* But the adoption of AI by many companies seems not at all incidental. So it is fair to ask how the introduction of AI is going to affect core Quality functions. How will AI change—say—auditing?

It's a good question, and there have been a couple of attempts to answer it in the past year. Last October, Elisabeth Thaller and Jorge Bravo CarreƱo wrote an article "When the QMS Thinks for Itself" for Quality Progress. Then just last month the ISO 9001 Auditing Practices Group published a formal guidance paper on "Auditing a Quality Management System that uses Artificial Intelligence (AI) systems." It is perhaps not surprising that the article and the guidance paper repeat the same basic points,** and they are pretty much the same points you would make if someone asked you about the same topic. Nothing here is unexpected.  

What makes this consistency possible is that ISO 9001 is structured to apply to any kind of work. So the basic questions are always the same. What are your processes? What are your risks? What are your training requirements? And are you getting the results that you want?

The article in Quality Progress identifies five main thoughts to keep in mind when planning an audit in an organization that uses AI, and the sugges­tions in the APG guidelines can almost be mapped directly onto the same list. The five main thoughts are these:

  1. Identify where the organization is actually using AI. ("AI may be hiding in plain sight.")
  2. Use the process approach to identify inputs and outputs in the normal way. ("The process approach still applies—even when the process thinks for itself.")
  3. Identify risks. ("Risk-based thinking isn’t optional.")
  4. Identify training needs, responsibilities, and authorities. ("Competence and oversight are essential.")
  5. Identify results and look for objective evidence. ("Focus on evidence, not the shine.")

And while I'm certainly not going to quote all ten pages of the guidance document in a blog post—you can download the whole thing for free by using the link above—let me quote just a couple of items that fit under each of these five headers, to show you the overlap.

Identify where the organization is actually using AI.

  • When preparing the audit, the auditor should:
    • Identify whether the organization uses any AI system type for any of its processes within the scope of the QMS.
    • Determine the specific function(s) of these processes within the QMS, and
    • Recognize assigned responsibilities in relation to these processes. [§2, p.4]

    Use the process approach.

    • Where AI system(s) are used within operational processes, has the organization ensured that their use supports the achievement of intended results?
    • Are such processes monitored, and are changes related to the use of AI system(s) controlled?
    • How does the organization ensure the consistency, correctness, and reliability of AI outputs? [§3, cl. 8, p.7]

    Identify risks.

    • Has the organization considered the implications of the use of AI system(s) when determining its internal and external issues?
    • Has the organization identified statutory and regulatory requirements in relation to the use of AI systems, such as those related to data privacy and information security? [§3, cl. 4, p.4-5]
    • Has the organization determined and addressed risks and opportunities related to the use of AI system(s)? [§3, cl. 6, p.6]

    Identify training needs, responsibilities, and authorities.

    • Are adequate resources available to maintain and update AI system(s), including access to technical expertise, data quality management, and cybersecurity support?
    • Is the organization able to demonstrate that personnel who use, manage, or oversee AI system(s) meet the determined competence requirements? [§3, cl. 7, p.6]
    • Does top management take accountability for ensuring that the AI system(s) within the QMS support its effectiveness and achievement of intended results? [§3, cl. 5, p.5]

    Identify results and look for objective evidence.

    • Are customer satisfaction trends evaluated for potential impacts arising from AI-enabled processes or interactions?
    • Is internal auditing addressing the effectiveness of AI-influenced processes?
    • Is information related to the performance and effectiveness of AI systems an input into management review? [§3, cl. 9, p.8]
    • Are nonconformities or complaints involving AI system outputs recorded, analyzed, and addressed? [§3, cl. 10, p.8]

    There shouldn't be anything shocking in this list. One of the strengths of ISO 9001 is precisely its flexibility. And that flexibility comes from seeing all kinds of work through a lens that highlights certain common features. AI is just one more topic viewed through the very same lens. This is why I said above that you would have come up with the same points if you had taken the time to work through the standard in detail.

    But most of us are too busy with our day jobs to do that, so it's convenient that the Auditing Practices Group has done it for us. The next time you have to audit an organization that uses AI in its management system, check out the guidelines document to see if any of the suggestions help you. 



    __________

    * In this post, I asked what the word “Quality” really means in the context of AI. In this one, I discussed how AI tools might be able to enhance training records. 

    ** Partly this outcome is unsurprising because the same person—Elisabeth Thaller—was the lead writer for the article and the chair of the committee that wrote the guidance document. Chris Paris published an article taking the APG committee to task because most of them have no special background in AI. But it appears from her LinkedIn profile that Thaller does have at least some background.        

               

    Thursday, March 5, 2026

    Gatekeeping

    Last week I wrote a post for this blog, and—as usual—tried to publicize it in LinkedIn. But I mostly failed. My very first post about the article (in my own timeline) went through well enough; but when I tried to cross post in special-interest groups, I invariably got the message, "Oops - we were unable to complete your request. Please try again later." And "trying again later" produced the same result. It was odd.

    LinkedIn Customer Support has hitherto offered a number of theories for what happened, none of which match the facts of the case. (Perhaps I explained the problem in a confusing way at first; but as we have emailed back and forth, I have tried to make the picture clearer.) But I have been able to post other notifications since then. So it is clear to me that my account has not been blocked, and I have not violated some overall limit on the number of posts. As far as I can tell, the only plausible explanation left is that my post triggered a content filter. The filter must have determined my post to be spam, not because there were too many copies of it (There weren't.) but somehow because of what I said. In other words, I must have run afoul of some gatekeeping protocol.     

    Gatekeeping isn't necessarily a bad thing, though in this case it proved inconvenient for me. In essence, gatekeeping is a form of preventive action. For example, many large companies have policies that require all their suppliers to be certified to ISO 9001. This is a gatekeeping requirement, and it is intended to weed out certain problems before they start. We have discussed at length that certification to ISO 9001 doesn't necessarily mean a company does good work. But it does mean that they have a way to handle complaints, and it does mean that they will work next week more or less the same way they worked this week. Those two reassurances—just by themselves—are huge. 

    By the same token, if you determine that a certain sort of message interferes with the purposes of your online community, it's only natural to take action in advance to prevent that kind of message from ever being posted. I assume that the LinkedIn filters probably thought my post was some kind of advertisement—maybe even an advertisement for some kind of artificial intelligence tool. While advertising is allowed on LinkedIn, it has to go through a certain procedure before it is posted. It would not surprise me to learn that LinkedIn has tools—maybe even AI tools, ironically—that screen individual posts to look for advertisements that have been disguised as personal notifications. It wouldn't even surprise me if someone showed me a list of the features that this tool looks for, the features that identify a probable advertisement-in-hiding, and it turned out that by a simple accident I had written some of those features into my post—so that the gatekeeping tool rejected it.

    It wouldn't surprise me. But at this point that's pure speculation. I have seen no such list yet, and LinkedIn Customer Service has not confirmed that this is the problem. I am just guessing.

    But let's take this line of thought one step farther. We all know that if you are going to derive preventive actions or lessons learned from a problem investigation, the actions or the lessons have to be based in some kind of causal way on the things you learned during the investigation. In the same way, if you decide to gatekeep your social media community to keep out AI-based advertising (or whatever it turns out to be), you should make sure that this is what you are really doing in practice. Look at your algorithm closely, and review it from time to time. Also, if someone's posts get caught by the algorithm and he complains, that might mean he's a human being (rather than an AI) so you might be able to use that information to refine your algorithm even more. 

    Meanwhile I'll try to write notifications that don't sound like ad copy. That's a preventive action on my part, to avoid this circumstance arising again.   

               

    Thursday, February 26, 2026

    Training systems and artificial intelligence

    A few days ago, I got a message from a man named Momchil Penev, asking what I think about training systems and their vulnerabilities? It was an interesting question, and I had to tell him that I haven't thought about the subject in depth. Then he had a few more questions, and we started a conversation that I'll explain in some detail below. But in the meantime I started to wonder ... why haven't I spent more time thinking about training systems?

    How important is training, really?

    We all know that training is important to Quality: if people don't know how to do a job, they can't be expected to do it well. It even shows up in ISO 9001—clause 7.2 (c) says explicitly that if your people don't already know how to do the work, you have to make sure they learn. But as an auditor, I've never found that checking training records contributes a lot to the overall outcome of the audit, other than by checking the box for clause 7.2. If I'm going to find problems in the organization, it seems like they are always somewhere else.

    But why? I think there are three reasons:

    • Training records usually consist of a list of classes (for office staff) or a list of operations (for manufacturing and logistics personnel), each one signed off by a supervisor or instructor. Often there is no indication what the signature actually means, so all I can check is whether the paperwork is complete. (Remember this point, because we'll come back to it.)
    • Most of the organizations I've audited are fairly small, or else they are small units inside large companies. It's hard for incompetence to hide in small companies, because people largely know what their neighbors are working on. I've been more likely to find that an employee's training record shows no signature for a skill that he knows very well, than to find a true gap in some mission-critical competency.
    • On the whole, when people fail to follow established procedures, I haven't found that ignorance is the root cause. There's always something else going on. At least in my experience, the employees who take shortcuts that end in disaster know what they are supposed to do, but they choose not to do it.   

    Sometimes training is not the problem!

    This last point is important. Remember Alaska Airlines flight 1282, that lost a door plug shortly after takeoff on January 5, 2024? The best explanation I've ever read of the root causes behind this accident comes from an anonymous account posted to the Internet by a current Boeing employee. You can find it here

    • This account makes the point that the door plug blew out because four bolts to hold it in place were missing.
    • The bolts were missing because they had been removed in order to fix a different problem with the door plug.
    • Then after that problem was fixed the bolts were never replaced because the workers forgot about them. 
    • But if they had followed the established repair procedures they could not have forgotten. The only way they could possibly forget was by taking shortcuts.
    • Why did they take a shortcut? Not because they were ignorant of the correct procedure, but because the procedure was too cumbersome and time-consuming. And they were already under pressure from management to hurry up.

    Or consider a story that I tell in this post from a few years ago. In a factory that I represented, an external auditor wrote up one of our manufacturing personnel because he was running a plating bath and the liquid was at the wrong temperature. The procedure told him that if anything was wrong he should shut down the line and call the responsible manufacturing engineer (who was on vacation at the time). But in fact this employee had decades of experience with plating. He knew that he could compensate for the wrong temperature by making minor changes to the chemical mix in the bath. The end product was indistinguishable from one made according to the instructions, and this way we could ship it to the customer on time rather than delaying the order two weeks until the engineer returned. We got an audit finding, but from the customer's perspective his solution was perfect.* In this case, the reason the employee failed to follow the written procedure was that he was confronting an unforeseen situation, and had the deep knowledge to improvise a solution.

    This is why I say that in my experience, ignorance of a procedure hasn't generally been the root cause of failure to comply. By the same token, when I am coaching a team through an 8D report to resolve some failure, I don't accept "Retrain employees" as a corrective action. There has got to be something more substantive behind the problem.  

    But training systems can still be more robust

    All that having been said, training systems are still necessary. And like anything else, they can always be improved. This brings me back to my conversation with Momchil Penev, who wants to use artificial intelligence technology to improve them.

    Penev is the founder of skillia.AI, and his idea is an interesting one. It's a bit of a right-angle turn in our discussion, but let me take a minute to explain what he is doing.


    As I say, it's an interesting idea. In general I'm something of an AI-skeptic; and as noted above, training by itself won't solve all Quality problems. But it still has to be done, and the records still have to be up to date! If Penev can find a problem and then solve it, that's how Progress happens. I wish him well. 


    POSTSCRIPT: During our conversation, Penev told me that he is eager to get feedback from people in the Quality business about the performance and limitations of his tool. This means you! So while I have no financial stake in Penev's success or failure, I told him I would post his Contact page. You can reach him at https://skillia.ai/contact. Tell him what you think of his ideas, or ask how you can test the tool for free. 

    __________

    * For further discussion of this finding, see the referenced post.     

     

          

    Thursday, July 31, 2025

    Does AI have a Quality problem?

    We've all seen articles about the incredible power and potential of Artificial Intelligence (AI). Whole industries are being restructured to make use of AI's capabilities. One article from last year—that I chose almost at random—lists the use cases for Large Lanuage Models (LLMs) as follows: 

    • Coding: "LLMs are employed in coding tasks, where they assist developers by generating code snippets or providing explanations for programming concepts."
    • Content generation: "They excel in creative writing and automated content creation. LLMs can produce human-like text for various purposes, from generating news articles to crafting marketing copy." [Gosh, are they any good at writing niche products like Quality blogs? šŸ˜€ Just kidding, of course!]
    • Content summarization: "LLMs excel in summarizing lengthy text content, extracting key information, and providing concise summaries."
    • Language translation: "LLMs have a pivotal role in machine translation. They can break down language barriers by providing more accurate and context-aware translations between languages."
    • Information retrieval: "LLMs are indispensable for information retrieval tasks. They can swiftly sift through extensive text corpora to retrieve relevant information, making them vital for search engines and recommendation systems."

    And so on. The article lists eight more use cases before summarizing with a list of half a dozen general benefits of LLMs. (I found myself wanting to ask if the author has an LLM in the family, perhaps as a favorite cousin or an in-law.) In short, LLMs can do quite a lot.

    But LLMs hallucinate! 

    We are discovering, though, that it is not safe to rely on LLMs for an accurate description of what is out there. When LLMs summarize content or retrieve information, sometimes they report things that aren't true. The first time I saw such a story, it was in this post from LinkedIn back in 2023, where Marcus Hutchins posted a conversation he had with the Bing AI chatbot. The bot claimed that it was still 2022, insisting "I know the date because I have access to the Internet and the World Clock"—even though it was verifiably already 2023!

    Then more stories started rolling in. To my mind the most dramatic has been the recent legal case SHAHID v. ESAAM (2025), Docket No: A25A0196, decided on June 30, 2025 by the Court of Appeals of Georgia. The summary description of this case makes for delightful reading, and I enclose a selection below in an extended footnote.* But the gist is that one party's pleading must have been generated by an LLM tool. No human lawyer could have written it. The pleading rested almost entirely on bogus case law: either cases that never happened, or cases that had no relation to the point at stake. This is the kind of mistake that junior paralegals get fired for. Even worse, the initial trial court accepted it without blinking. The bogus citations were caught only by the Court of Appeals.

    So employing LLMs comes with a risk. You can't just blindly trust whatever they tell you without cross-checking it, because they fabricate content so effortlessly. Looking back at that list of use cases at the top of this post, I have to qualify the claim that you can use them in writing or summarizing: maybe LLMs can suggest an interesting idea you didn't think of before, but they can't do your work for you. Some people, though (like Sufyan Esaam's attorney) want to use them for just that.

    It's a problem.

    What about Quality?

    But is it a Quality problem? Here the answer is not so clear, because it depends how exactly you define Quality. You remember that I prefer to define Quality as "getting what you want"; and in that sense—especially if "you" means the end-user of the AI tools—then AI hallucinations constitute a big Quality problem. When AI hallucinates a false answer to my question, I'm not getting what I want.

    But there is another definition, which says that "Quality is conformance to requirements." And with that definition the situation is rather different ... because the LLM programs are doing exactly what they have been told to do! 

    Jason Bell of Digitalis.io made this argument in a recent LinkedIn post. The point is that the LLM tool is not programmed to see what's really there. It is not programmed to perceive reality, and it is not programmed to tell the truth. Its only programming is to say something that sounds good, subject to certain parameters that define what it takes for something to "sound good." But perceiving reality and telling the truth are never part of that definition, because AI has no mechanism or equipment to allow it to carry out those tasks.

    In a sense, then, the problem is not with AI itself, but with user expectations. It's like if I use a hammer to comb my hair: the results are pretty sketchy, but it's not the hammer's fault.  

    Of course I have no idea how long AI will be a big deal, or how much of an impact it will have on our work and our lives. But as long as it is here, it will be useful for us to be clear on its capabilities and limitations, so that we can distinguish reality from science fiction. In its current form, AI has no cognitive component, and therefore cannot observe reality or distinguish truth from falsehood. But it is very good at sifting through piles of words according to defined rules.

    And honestly? That's enough for now. Let's use it for the things it really can do, and not try to make it comb our hair or understand the world. After all, if AI ever develops a cognitive component, that addition will doubtless bring new problems of its own. 

    HAL 9000, of course.

    __________

    [Emphasis is mine, in all cases.] 

    After the trial court entered a final judgment and decree of divorce, Nimat Shahid (“Wife”) filed a petition to reopen the case and set aside the final judgment, arguing that service by publication was improper. The trial court denied the motion, using an order that relied upon non-existent case law. For the reasons discussed below, we vacate the order and remand for the trial court to hold a new hearing on Wife's petition. We also levy a frivolous motion penalty against Diana Lynch, the attorney for Appellee Sufyan Esaam (“Husband”)....

    Wife points out in her brief that the trial court relied on two fictitious cases in its order denying her petition, and she argues that the order is therefore, “void on its face.”

    In his Appellee's Brief, Husband does not respond to Wife's assertion that the trial court's order relied on bogus case law. Husband's attorney, Diana Lynch, relies on four cases in this division, two of which appear to be fictitious, possibly “hallucinations” made up by generative-artificial intelligence (“AI”), and the other two have nothing to do with the proposition stated in the Brief.

    Undeterred by Wife's argument that the order (which appears to have been prepared by Husband's attorney, Diana Lynch) is “void on its face” because it relies on two non-existent cases, Husband cites to 11 additional cites in response that are either hallucinated or have nothing to do with the propositions for which they are cited. Appellee's Brief further adds insult to injury by requesting “Attorney's Fees on Appeal” and supports this “request” with one of the new hallucinated cases.

    We are troubled by the citation of bogus cases in the trial court's order. As the reviewing court, we make no findings of fact as to how this impropriety occurred, observing only that the order purports to have been prepared by Husband's attorney, Diana Lynch. We further note that Lynch had cited the two fictitious cases that made it into the trial court's order in Husband's response to the petition to reopen, and she cited additional fake cases both in that Response and in the Appellee's Brief filed in this Court.

    As noted above, the irregularities in these filings suggest that they were drafted using generative AI....

    Five laws of administration

    It's the last week of the year, so let's end on a light note. Here are five general principles that I've picked up from working ...