Talk to an engineer
Whether you have a full brief or only a rough idea, send it over. Every inquiry goes straight to our team.
Send us a message
Data Preparation: The Work That Decides Your Model's Quality
9 Oct 20262 min read
Labelling, cleaning, splitting and balancing: why the least glamorous part of an AI project drives the results.
Data Preparation: The Work That Decides Your Model's Quality
Model architectures are compared endlessly; the data feeding them is discussed far less, even though it usually decides the outcome. A well-prepared dataset with a plain model reliably beats a sophisticated model fed with noise.
Collect with a purpose
Start from the decision the model will support and work backwards to the examples that demonstrate it. Gather from where the work already happens: tickets, orders, documents, records. Then check the practical questions: volume, quality, staleness, and whether the data may legitimately be used for this purpose.
Labels are an argument
Where humans label the data, the instructions matter more than the labeller. Write the rules, resolve the disagreements, and record which cases were ambiguous. Inter-annotator agreement is not bureaucracy; it is a measurement of whether your task is actually well defined. If two people cannot agree, the model cannot be expected to.
Clean in ways you can defend
Duplicates, missing values, outliers, inconsistent formats, text from one source that is twice as long as another. Every cleaning step should be reproducible, recorded and reversible, so the dataset that produced a result can be rebuilt later.
Split honestly
Training, validation and test data divided so nothing leaks between them, and divided by the thing you care about: by customer, by time, by document, not by random row. A random split that puts near-identical records on both sides produces impressive numbers that collapse in production.
Watch the imbalance
Most real problems are imbalanced: the rare case is the interesting one. Accuracy on a dataset where the majority class dominates tells you nothing about the case the project exists to catch.
Keep a baseline
Compare against how the task is done today, on held-out data. If the model cannot beat the current process, the honest answer is that it is not ready, whatever the loss curve says.
How we work
AI and machine learning at Black Origin IT spends real time here, before architecture, because preparation is where the project is won.
Sitting on data with questions in it? Tell us about your project and we will see what it can support.
