Conversational AI with a chatbot is great for drafting emails or debugging code, but it’s less ideal for real-time application middleware. If you’re trying to inspect a financial transaction for potential fraud in the middle of a checkout loop, you don’t need an LLM to write you an essay about why a credit card transaction looks suspicious – you just need a probability score, and you need it as fast as possible.
To this end, TypeSafe AI recently released Jev, a cloud-based service that uses modified generative AI models for classification tasks. Naturally, the open weight model community has responded with similar, publicly accessible models that you can run yourself. Instead of generating natural language, these AI large language models bypass text generation entirely. They read the context you provide them with, evaluate all your typed questions in a single parallel pass, and report back the probabilities.
The advent of this new LLM-based classification technique seemed to me to be, if nothing else, a great opportunity to demonstrate how to use Canonical’s open source MLOps stack to build a realtime transaction fraud detection system, based on an open weight, Jev-like model – the “system-one-qwen3.5-4b-scorer” model.
Sounds like fun? Good. Let’s start.
Open source MLOps
Canonical’s open source MLOps stack is pretty comprehensive, with integrated support for Kubeflow, Katib, Feast, MLFlow and KServe for doing everything from pre-training, to fine tuning and reinforcement learning, with a set of structured tools to help with the full project lifecycle from early research to production.
But for this project, we’ll just be using KServe. KServe is a Kubernetes based inference server control plane. It helps you to deploy AI models – including large language models – on Kubernetes, so that you can benefit from the resilience and automation that Kubernetes offers for production deployments. Before we begin in earnest, make sure you you can meet the prerequisites:
- You’ll need to be running at least Ubuntu 24.04 LTS on your computer.
- Your computer needs to have at least 32GB of RAM and a modern CPU to run everything comfortably.
- I initially used the integrated graphics processor on my laptop’s CPU to accelerate the model’s processing, but to make this lab project more accessible, we will use your computer’s CPU. The model will just run slower.
- A stable, high speed internet connection for downloads.
- Some familiarity with Linux commands so that you can follow along.
Install KServe and set up the lab project
First up, let’s prepare the system. Open up your terminal and make sure you’ve got the base software installed. Run the following commands:
sudo snap install microk8s --channel 1.34-strict/stable
sudo microk8s enable storage dns
sudo microk8s enable metallb 10.64.140.43-10.64.140.49
sudo snap install juju --channel 3.6/stable
sudo microk8s.config | juju add-k8s my-k8s
sudo snap install terraform
sudo apt install git -y
git clone https://github.com/canonical/charmed-kubeflow-solutions.git
pushd charmed-kubeflow-solutions
git checkout track/1.11
pushd terraform/products/kubeflowCongratulations, you’ve just installed a one-node, super-compact Kubernetes cluster running Canonical’s MicroK8s, with Canonical’s operations management engine Juju, and with Terraform.
Now let’s launch Charmed Kubeflow. We’re only going to enable Kubeflow and KServe, so we’ll disable all the other modules like Katib, MLFlow, TensorBoard, federated login and the Canonical Observability Stack. You can always enable them if you want, but they aren’t needed for the purposes of this demo.
DEX_USERNAME=username
DEX_PASSWORD=password
terraform init
terraform apply -auto-approve -var dex_static_username="${DEX_USERNAME}" -var dex_static_password="${DEX_PASSWORD}" -var enable_feast=false -var enable_katib=false -var enable_kserve=true -var enable_training_v2=false -var enable_training_v1=false -var enable_observability=false -var enable_tensorboard=false
popd
popdGive it 10-15 minutes for things to come up on the Kubernetes node. You can monitor how close things are getting to completion with the command `juju status -m kubeflow`. When everything is saying “idle/active”, we’re ready to go further with this lab project.
I’ve prepared the scripts needed to successfully complete this project and put them up on GitHub. The next step will be to clone the git repository to your local system. Run the following command:
git clone https://github.com/grobbie/realtime-transaction-fraud-detection-with-an-llm.git
pushd realtime-transaction-fraud-detection-with-an-llmNow let’s install the development tools that we’re going to need. Make sure not to skip these steps.
sudo apt update && apt install -y python3-venv python3-pip
python3 -m venv .venv
source .venv/bin/activate
.venv/bin/pip3 install -r requirements.txtModel crush
The model we’re using isn’t the biggest, but it’s still big enough to crash a laptop – or rather, to crash the Kubernetes pod it’s going to run inside. So there are a couple of tricks we’re going to use to make it work much more efficiently on your computer, even without a high power GPU.
The model we’re using is actually a derivative model using a LoRA adapter architecture. Low rank adaptation (LoRA) is a model fine-tuning technique that allows you to adapt a large language model to a specific task, without modifying the original model at all. However, for our purposes this approach adds additional overheads – every request has to pass through both the base model and the adapter paths simultaneously. This is good if you want to hot-swap model specialisms on a base model in real-time. But we just want a compact, efficient model that’s going to do classification tasks, so we’ll merge the LoRA weights into the base model.
The other trick we’re going to use is Activation-aware Weight Quantization (AWQ). This is a 4-bit post-training quantization technique that compresses the model with minimal loss of accuracy. With AWQ, we can squash the amount of VRAM the model needs down by up to 75%. While this technique is really designed and optimized for accelerating models that are running on GPUs, it does also work on the CPU. So for our Qwen3.5-4b derivative model, we should be able to get the model’s VRAM needs down from about 18GB all-in to around 5GB. Not too bad and definitely enough of a squeeze for you to be able to reliably run it on a laptop or similar without crashing, and yielding decent performance.
Run the following command to download a calibration dataset and prep it, and then merge the model weights and quantize:
.venv/bin/python3 quantize_awq.pyThis downloads the model weights from Hugging Face Hub, runs the merge and quantization, and produces a `./system-one-4bit-awq` directory containing all the necessary model artifacts. The task could take a while to complete – expect it to take around 6 hours on a laptop because of the calibration that we need to do during quantization.
Run the model on vLLM
Now we’ve got the model, we need a way to run it. We’ll run it on the industry standard vLLM model server, using KServe as a means to set everything up nicely on Kubernetes.
To deploy the vLLM serving pod and service, you can use the manifests from my git repository. Before you run the commands, be sure to edit the file `kserve_deployment_cpu.yaml` to adjust the path to the model. Make sure the path is set to wherever the `system-one-4bit-awq` directory is on your system.
microk8s.kubectl create namespace modelserving
microk8s.kubectl apply -f kserve_deployment_cpu.yaml
microk8s.kubectl apply -f kserve_service_lb.yamlIt’ll take a fair bit of time for vLLM to start up, load the weights into memory, and get ready to serve on your CPU. You might be waiting as long as 5 minutes for this depending on the specs of your system. Obviously, it almost goes without saying that if we were doing this for real we’d use GPU acceleration and the whole experience would be a lot more performant.
To verify the pod is running, you can try the following command:
microk8s.kubectl get inferenceservice -n modelservingEventually, if all goes well the output will show that the health checks are good, and you’ll be ready to try out some classification tasks with your new specialized artificial brain.
Run some predictions
Obviously the goal of all of this, besides learning, is to run some transaction fraud detection with an LLM. So we want to try that, right? Of course. Let’s try that now and see if it works.
First we’ll try with a terminal command, and then we can try with a little web app I put together. To make sure the model is running accurately, we’re going to send it a web request using the `curl` command in the terminal, and see how it responds. We can use this one-liner command to find the right IP address for the vLLM model server:
KSERVE_IP=$(kubectl get services -n modelserving | grep system-one | grep private | awk '{ print $3 }')Now let’s send a request to the vLLM endpoint and see what we get back:
curl -s -X POST http://${KSERVE_IP}:8080/v1/completions
-H "Content-Type: application/json"
-d '{
"model": "/model",
"prompt": "<|im_start|>usernSystem One Fraud Risk Decision Query:nnTarget Transaction Profile:nTRANSACTION PROFILE:n- Type: TRANSFERn- Transaction Amount: $100,000.00n- Sender Initial Balance: $100,000.00n- Sender Post-Transaction Balance: $0.00 (Delta: $100,000.00)n- Recipient Initial Balance: $0.00n- Recipient Post-Transaction Balance: $100,000.00 (Delta: $100,000.00)nnQuestion: What is the fraud risk classification for this transaction?nOptions: 0: LOW_RISK, 1: MEDIUM_RISK, 2: HIGH_RISK_FRAUD<|im_end|>n<|im_start|>assistantnAnswer: ",
"max_tokens": 1,
"temperature": 0.0,
"logprobs": 3
}' | .venv/bin/python3 -m json.toolThe request contains the context – a transaction record, showing the transaction type, amount, the sender’s initial balance, the sender’s balance after the transaction, the recipient’s initial balance, and the recipient’s post-transaction balance. It also contains a structured JSON object containing the question: “What is the fraud risk classification for this transaction?”, and it contains the options that the model can use – “LOW_RISK”, “MEDIUM_RISK”, and “HIGH_RISK_FRAUD”. Each option has a short description. You could actually ask a few of these structured questions, and they should all be processed in parallel by the model.
If everything goes correctly, then when you run the command you’ll get some output back pretty quickly. It should look something like the listing below which is a response giving the model’s prediction.
{
"id": "cmpl-84ef1f97a3cc8d37",
"object": "text_completion",
"created": 1789839558,
"model": "/model",
"choices": [
{
"index": 0,
"text": "2",
"logprobs": {
"text_offset": [0],
"token_logprobs": [-1.5500739812850952],
"tokens": ["2"],
"top_logprobs": [
{
"2": -1.5500739812850952,
"1": -1.7219489812850952,
"0": -2.0422616004943848,
}
]
},
"finish_reason": "length"
}
],
"usage": {
"prompt_tokens": 182,
"total_tokens": 183,
"completion_tokens": 1
}
}It can be a bit tricky to interpret, but basically the bits to look for are the value of `text` and `top_logprobs`. “text” indicates the predicted classification (“text”: “2”). The model selected option token “2”, which corresponds to HIGH_RISK_FRAUD.
Looking back at the transaction record we sent, the transaction was a wire transfer of $100,000 to an empty bank account that completely drained the sender’s account. Definitely a risky transaction highly likely to be fraudulent. So if your model responded with “text”: “2”, then it got it right.
Not as much fun using the terminal though, so let’s start up that web app I promised you, which you’ll find in my repo. Start it by running the following command:
.venv/bin/python app.pyThe interactive web UI should be accessible at http://localhost:8000/ for you to play with.
Take it further
This has just been a lab project, but you can definitely take this further. Open weight models and open source software enable you to run AI behind the firewall in your own, self-governed environment – without needing to use cloud services. And with some optimizations the investment involved to run this technology doesn’t need to be out of reach – especially for enterprise use cases like this one. Of course, the whole process from quantization of the model to running the model could be accelerated with only a modest investment in GPUs; but if you want to run this kind of workload in production with low milliseconds response times across a huge number of concurrent real time transactions, then you might need a more substantial setup.
In any case, Canonical’s MLOps stack is there to support your use case.
- Discover the full capabilities of Charmed Kubeflow
- Want us to manage the infrastructure so you can focus on the models? Sign up for Managed Kubeflow on Azure
Discover more from Ubuntu-Server.com
Subscribe to get the latest posts sent to your email.

