Why I’m Testing Local LLMs This Summer
(And Why It Matters for HR & Payroll)
Earlier this year, I listened to an episode of the Baseline podcast that I found so interesting that I listened to it several times (you should too!). Podcast host Ian Smith interviewed Jonny Williams, the Chief Digital Adviser for the UK Public Sector at Red Hat. He helps the government to figure out what to actually do with AI, as opposed to what they think they should do with it. This conversation wasn’t about HR or payroll. I don’t think they mentioned it even once. But I’ve been thinking about the topics they discussed ever since, and I can easily translate them to HR cases.
This blog is the first in a series I’m writing this summer as I run my own experiment: testing local LLMs, open source language models to understand if and what they can realistically do for HR. I want to start by explaining why I’m doing this, because the reasons matter as much as the results (maybe even more!).
We haven’t found the right story yet
One of the things Williams said that stayed with me was that organizations are failing to connect AI to outcomes that people can actually see. The technology exists and the investments continue at a rapid pace. But the story isn’t there yet: a clear, honest account of what AI is actually doing and for whom. How it makes us and our lives better. And when people don’t see the value, they disengage. He illustrates that with an example in government, where significant money is being invested in AI, while people’s most visible experience is a chatbot that they really don’t want to use. And that they actively try to bypass by typing “I want to speak to a human.”
I see the same pattern in HR. In the past two decades, we’ve experienced several waves of new technology. They were going to make our work-life better. But in the end, what we’re left with doesn’t always match the early promise. If an HR professional’s first real experience with an AI tool is that it doesn’t understand their organizational context, or produces outputs they can’t trust, or creates more work in checking the outcome than it saves, that shapes how they think about AI. And those first impressions are hard to overcome. When we want to get real value out of AI, we must be honest in telling the story of what it can and can’t do.
AI should take on the work that drains people, not the work that defines them
Williams referred to a phrase I often use in keynotes: “I want AI to wash the dishes and do the laundry.” (I really do!) AI should take on the administrative burden, the repetitive tasks, the things that take up time and energy and do not need the judgment and experience that HR professionals have spent years developing. He is critical of organizations that focus AI on replacing creative and skilled work, because that misses the point entirely.
For me, that means the question in HR is not whether AI can replicate what a good HR professional does. Instead, we should find out if it can take enough of the weight off their shoulders so they can spend time on the work that actually requires their skills and drives value, for the organization and its employees. That’s what I want to test with these local LLMs. Things like a first draft of a policy or accurately summarizing long documents like collective labor agreements. These are tasks where a reliable first pass can save time, even if an HR professional still needs to review and refine it. I might move to some of the more complex tasks later, but first things first.
Technology decisions should come after the people and process questions
Williams also makes a point that I’ve spoken about many times: the decision about which technology you use comes after you understand the problem, the people and the process, not before. Right now, in many organizations, it’s the other way. We are so in awe of AI that we are letting technologists drive AI decisions. The people who understand how HR and payroll actually works, from the operational complexity to the human sensitivity of the data and the regulatory constraints, are being consulted after the fact. Or are called laggards and luddites, even though their concerns are valid.
I find this genuinely worrying. Because in HR, getting this sequence wrong has real consequences. Not just for the project, but for the people whose data and working lives are involved. For organizations that might be found liable when there’s a breach in data privacy and security. My approach this summer is to start exactly where I think organizations should start: with specific HR use cases, a clear understanding of what data is involved and what the constraints are, and an honest assessment of what local LLMs can do within those boundaries. I will also compare the outcome to frontier models. But in my approach, the technology is the last question, not the first.
From one giant model to smaller, more specific local LLMs
There’s something else in the podcast conversation that I think is worth sharing. Williams draws a comparison to what happened in software development when organizations moved away from large, monolithic systems (one platform that did everything) towards smaller, specific tools that work together (best-of-breed). He thinks that AI is heading in the same direction.
Right now, most organizations are thinking about AI as one big model that knows everything and can be accessed through a single interface. We use the frontier models for every topic, from personal to business and they can handle it all. I’ve always wondered why these vendors are not building focused models for business. And like me, Williams is more interested in what comes next: smaller models, trained for specific purposes, that are accurate and trustworthy within their domain.
He puts it like this: a model advising on tax governance doesn’t need to know Roman poetry, just as a model supporting children’s education doesn’t need capabilities that serve far darker purposes. I think that same logic applies to HR. A model that knows employment law, HR policy language, and payroll compliance doesn’t need to be a general-purpose AI that’s been trained on everything. You would not use it to answer questions about car maintenance or beauty routines. It needs to be focused only on the context, like HR, and be capable about the things that matter in that context.
That should also influence cost. The frontier models charge per token. You can think of that as a fee per word processed. When you are running multiple queries at the same time, that adds up quickly, especially now that the vendors are increasing token pricing and/or throttling use. Running local LLMs on your own hardware could be less expensive: you need the machine itself plus training and support. That’s part of what makes local, smaller models interesting to me. Not just as a focus question, but from a sustainable cost perspective. But I could wrong in underestimating the effort to run and maintain local models.
Why local models, and why now
I’ve bought a separate laptop for this experiment so I can use a clean machine that is not connected to my work files. I want to test this without any risk to the data I work with. I’m mostly curious about the capabilities of open source, local LLMs. But it’s about more than that:
- It’s about explainability: I want to be able to understand what a model is actually doing, rather than accepting outputs from something I can’t inspect.
- It’s about data sovereignty: In HR you must know where data goes, who can see it, and what happens after you submit a query or manipulate it.
- It’s about ownership: Building HR processes on top of a model that is constantly changing is a risk worth thinking carefully about.
- It’s about cost predictability: I’ve written before that the pricing structures around the big frontier AI models will only go up.
Being able to control what you use matters because in HR, privacy and security aren’t abstract concerns. HR handles some of the most sensitive data: Salary information, performance records, health disclosures, family circumstances that affect working arrangements. Using AI tools in HR without understanding what happens to that data when you upload it, is a technology risk as well as a professional liability. Local models, that run on your own hardware in your corporate network and don’t send data to a third-party server, offer a way to manage that. I want to understand, practically, what that actually looks like and what the trade-offs are.
Over the coming weeks, I’ll share what I find: which local LLMs I am using, what these models can do well, where they fall short, and what it means for HR teams who are thinking seriously through their AI strategy. I’ll report back what I find out, even when the results are disappointing. Because when new technology comes along, we should keep an open mind and we owe it to ourselves to find out how it can help us or not.
