Data Preparation: The Work That Decides Your Model's Quality

9 Oct 20262 min read

Labelling, cleaning, splitting and balancing: why the least glamorous part of an AI project drives the results.

Data Preparation: The Work That Decides Your Model's Quality

Model architectures are compared endlessly; the data feeding them is discussed far less, even though it usually decides the outcome. A well-prepared dataset with a plain model reliably beats a sophisticated model fed with noise.

Collect with a purpose

Start from the decision the model will support and work backwards to the examples that demonstrate it. Gather from where the work already happens: tickets, orders, documents, records. Then check the practical questions: volume, quality, staleness, and whether the data may legitimately be used for this purpose.

Labels are an argument

Where humans label the data, the instructions matter more than the labeller. Write the rules, resolve the disagreements, and record which cases were ambiguous. Inter-annotator agreement is not bureaucracy; it is a measurement of whether your task is actually well defined. If two people cannot agree, the model cannot be expected to.

Clean in ways you can defend

Duplicates, missing values, outliers, inconsistent formats, text from one source that is twice as long as another. Every cleaning step should be reproducible, recorded and reversible, so the dataset that produced a result can be rebuilt later.

Split honestly

Training, validation and test data divided so nothing leaks between them, and divided by the thing you care about: by customer, by time, by document, not by random row. A random split that puts near-identical records on both sides produces impressive numbers that collapse in production.

Watch the imbalance

Most real problems are imbalanced: the rare case is the interesting one. Accuracy on a dataset where the majority class dominates tells you nothing about the case the project exists to catch.

Keep a baseline

Compare against how the task is done today, on held-out data. If the model cannot beat the current process, the honest answer is that it is not ready, whatever the loss curve says.

How we work

AI and machine learning at Black Origin IT spends real time here, before architecture, because preparation is where the project is won.

Sitting on data with questions in it? Tell us about your project and we will see what it can support.