Oracle AI & LLMQLoRA & Fine-Tuning

An Old DBA Learns AI: I Tried to Train an LLM for the First Time

Gökhan Dedeler
Gökhan Dedeler
Senior Principal Oracle DBA • Founder @ DBADoctor.com
Aug 22, 2026• 11 min read

I've been an Oracle DBA for many years.

I've spent most of my career working with databases, RAC, ASM, Exadata, Data Guard, performance tuning and troubleshooting.

Then I started experimenting with AI.

My first experiments were around embeddings, vector search and RAG.

This time, I wanted to try something different:

What happens if I actually try to train an LLM with my DBA knowledge?

Not just give the model information through RAG.

Not just retrieve information from a vector database.

Actually train the model.

So I opened a Google Colab notebook, connected a Tesla T4 GPU, loaded Qwen 2.5 7B and attempted my first QLoRA fine-tuning experiment.

And the result was not exactly what I expected.

The Goal

The idea was simple.

I wanted to take a general-purpose LLM and give it a small piece of DBA knowledge.

For the first experiment, I deliberately kept things extremely small.

I used one Oracle DBA question:

How do you troubleshoot ORA-01555 Snapshot Too Old?

And one short answer:

Check UNDO sizing and retention settings, long-running queries, and undo generation.

That's it. One question. One answer.

This wasn't intended to create a production DBA model. It was an experiment to understand what actually happens when you fine-tune an LLM.

The Hardware

I used Google Colab with a Tesla T4.

The environment reported:

PyTorch: 2.11.0+cu128

CUDA available: True

GPU: Tesla T4

VRAM: 14.56 GB

The GPU had 15 GB of physical memory, with approximately 14.56 GB available to PyTorch.

That was enough to experiment with a 7B model using quantization and QLoRA.

The Model

I used:

Qwen/Qwen2.5-7B-Instruct

The model has roughly 7.6 billion parameters.

A full fine-tuning run would be much more demanding than what I wanted to do on a Tesla T4. That's where QLoRA comes in.

What Is QLoRA?

The simple explanation:

Instead of updating all 7.6 billion parameters, we keep the original model largely frozen and train a much smaller set of additional parameters using LoRA.

The model was configured with:

r = 8

lora_alpha = 16

lora_dropout = 0.05

target modules:

q_proj, k_proj, v_proj, o_proj

And this was the interesting part. The training output showed:

trainable params:5,046,272
all params:7,620,662,784
trainable%:0.0662

So I wasn't training the entire 7.6 billion parameter model. I was training approximately 5 million parameters. That's only 0.0662% of the model.

For me, this was one of the most interesting moments of the experiment.

Preparing the Training Data

I started with a single training example:

training_data = [

  {

    "instruction": "How do you troubleshoot ORA-01555 Snapshot Too Old?",

    "output": "Check UNDO sizing and retention settings, long-running queries, and undo generation."

  }

]

I converted it into a Hugging Face Dataset:

Dataset({ features: ['instruction', 'output'], num_rows: 1 })

Then I converted the example into Qwen's chat format:

<|im_start|>system

You are Qwen, created by Alibaba Cloud. You are a helpful assistant.

<|im_end|>

<|im_start|>user

How do you troubleshoot ORA-01555 Snapshot Too Old?

<|im_end|>

<|im_start|>assistant

Check UNDO sizing and retention settings, long-running queries, and undo generation.

<|im_end|>

The final example contained 64 tokens. Then it was converted into input_ids, attention_mask, and labels.

At this point the data was ready for training.

My First Training Run

I configured a very small experiment:

Batch size: 1

Learning rate: 0.0002

Epochs: 3

FP16: True

Gradient steps: 1

Then I created the Hugging Face Trainer and ran train_result = trainer.train().

This was the first moment where the model was actually being trained.

And It Worked...

Technically. The training completed all three epochs. The training loss was:

StepTraining Loss
14.630938
24.137739
33.824971

So the loss went from 4.63 → 3.82. That was exciting to see.

My first thought was: "It worked. The model is learning." But then I tested the model. And that's where things became much more interesting.

The First Problem

My first post-training generation produced a strange repetitive output:

The problem. 2.4. 1. 19. 1. 1. 1. 1. 1. 1. 1...

Obviously, this wasn't a useful DBA answer. There were also warnings related to gradient checkpointing and use_cache.

So I switched the model to inference mode, disabled gradient checkpointing and tested it again. This time the output looked much more normal. The model produced a long Oracle troubleshooting answer.

But there was another problem: The answer wasn't necessarily correct.

Lower Loss Doesn't Mean Better DBA Knowledge

The model produced an Oracle-style troubleshooting answer, but some of its recommendations were questionable. For example, it focused heavily on locks and long-running transactions and suggested approaches that weren't necessarily appropriate for ORA-01555 troubleshooting.

So I decided not to stop there. I asked a different question:

"What are the main causes of ORA-01555 in Oracle Database?"

The model again produced a detailed Oracle-looking answer talking about long-running transactions, concurrency, row versioning, read consistency, and isolation levels.

But it also introduced questionable information, including a reference to a DBSNMP user's DB_VERSION_RETENTION parameter. That wasn't the kind of reliable DBA knowledge I was hoping to get.

The Final Test

I wanted to see whether the model could distinguish between two completely different Oracle errors. So I asked:

"What is the difference between ORA-01555 and ORA-00600?"

The model correctly described ORA-00600 as an internal Oracle error. But its explanation of ORA-01555 was wrong: It claimed that ORA-01555 was related to a table or index being altered, dropped or renamed during a transaction!

That was a clear factual problem.

The Model Had Been Trained — But It Hadn't Become a DBA

The technical training pipeline worked. The QLoRA adapters were trained. The loss decreased. The model changed its behavior. But that didn't mean it had learned reliable DBA knowledge.

My experiment was based on: One question + one answer. That's nowhere near enough data to teach an LLM a complex domain like Oracle Database administration.

A lower training loss does not automatically mean better domain accuracy.

What Did I Actually Learn?

This experiment changed the way I think about fine-tuning. Before this experiment, I was thinking about fine-tuning mainly as:

Give model DBA knowledge
↓
Train model
↓
Model knows DBA

But the real process is much closer to:

Training data
↓
Quality of examples
↓
Training configuration
↓
Fine-tuning
↓
Model behavior
↓
Evaluation
↓
Domain accuracy

The training data matters enormously. If the training dataset is tiny, incomplete or contains incorrect information, fine-tuning won't magically fix that. It can actually make the model behave differently without making it more trustworthy.

Fine-Tuning vs RAG

This experiment also made the difference between my previous RAG experiment and fine-tuning much clearer.

With RAG, the idea is roughly:

User question → Embedding → Vector Search → Retrieve relevant DBA knowledge → LLM → Answer

The knowledge lives outside the model.

With fine-tuning, we're changing the model's behavior through training:

Training examples → QLoRA / Fine-Tuning → Model → New behavior

That distinction became much more real to me after actually running the experiment.

Was the Experiment a Failure?

I don't think so. The technical experiment succeeded. I managed to:

  • load a 7B LLM on a Tesla T4
  • use quantization
  • configure QLoRA
  • train only 0.0662% of the model parameters
  • complete a training run
  • observe the training loss decreasing
  • generate responses after training
  • test the model on questions it had never seen

But the knowledge experiment was not successful. And that's actually more interesting, because I now have a much better understanding of what I need to do next.

What Would I Do Differently?

The obvious next step is to increase the quality and size of the training dataset.

Instead of 1 question + 1 answer, I would build a proper DBA dataset containing examples around:

• Oracle errors
• Performance troubleshooting
• AWR / ASH
• RAC
• ASM
• Data Guard
• Exadata
• RMAN
• SQL tuning
• Troubleshooting procedures
• Real-world DBA scenarios

And, most importantly, I would need to validate the answers before using them for training.

Then I could compare: Base model vs RAG vs Fine-Tuned model using the same evaluation questions. That would be a much more meaningful experiment.

The Complete Experiment

I wanted this experiment to be reproducible and transparent. So I published the complete Google Colab notebook, including the code, model loading, configuration, dataset preparation, training logs, warnings, outputs and the results — including the parts that didn't work as expected.

View the complete DBADoctor V2 training notebook on GitHub

I deliberately kept the logs. This isn't a polished AI demo. It's my first attempt at training an LLM, and I wanted to document what actually happened.

View Notebook on GitHub

Final Thoughts

I'm still an Oracle DBA learning AI. And that's probably the most interesting part of this journey.

My first experiment taught me about embeddings. My second experiment showed me how RAG can bring external knowledge to an LLM. This experiment taught me that fine-tuning is not simply "put knowledge into the model."

The model can train. The loss can go down. The output can change. And it can still give you the wrong answer.

So for now, my biggest takeaway is: Training the model is only half the job. Proving that the model learned the right thing is the other half.

And that's where my next experiment begins. DBADoctor V2 is still learning.

Next: A larger, properly validated DBA dataset — and another attempt. 🚀

Tech Stack

Google Colab
Tesla T4
Qwen/Qwen2.5-7B-Instruct
PyTorch
Transformers
PEFT & QLoRA
BitsAndBytes
Hugging Face Datasets & Python