Yarnin Peled YP Monogram Logo
Yarnin Peledponyapp.net
Enterprise AI Insight
July 20266 min read

Sort Your Data First, Before You Do Anything Else With AI

Why so many ambitious AI initiatives stall not in the algorithm, but in the archive.

#Data Governance#AI Readiness#Data Quality#Digital Transformation#Enterprise Architecture

Why do so many ambitious AI initiatives stall not in the algorithm, but in the archive?

And if artificial intelligence is only ever a mirror of the information it is given, what precisely are we asking it to reflect?

The prevailing enthusiasm for AI adoption tends to begin at the wrong end of the problem. Organisations convene workshops to brainstorm "AI ideas", commission proofs of concept, and evaluate vendors, all before confronting the far less glamorous question of whether their underlying data is fit to be read by a machine at all.

The discipline that actually determines success is quieter and considerably less exciting. It begins with workflow and process mapping: identifying where employees spend the most time on repetitive, rules-based tasks, and where sufficient, reliable data exists to support automation or augmentation.

But mapping the work is only the preliminary. The single most consequential activity an organisation can undertake before implementing AI is the restructuring of its foundational information.

An AI model is, in the end, a reflection of what it ingests. Feed it order, and it returns clarity. Feed it chaos, and it will not merely fail; it will amplify that chaos, reproducing every inconsistency, duplication, and contradiction at machine speed and with an authoritative tone that makes the errors harder to detect.

The uncomfortable truth is that most enterprises do not have a data problem they can see; they have a data problem they have learned to work around. Employees quietly reconcile conflicting figures, remember which spreadsheet is the "real" one, and know to ignore the folder marked final_v2_USE_THIS. An AI agent inherits none of that institutional intuition.

Section

Bringing Order to Databases and Information Sources

Disparate databases cannot meaningfully serve an AI model without strict standardisation. Three disciplines are non-negotiable. The first is schema alignment: translating the varying structures, field names, and data types that accumulate when systems are procured independently over many years. The second is source consolidation: identifying the single source of truth for each core business metric and systematically eliminating the redundant or conflicting sources that compete with it. The third is data cleaning: addressing missing variables, duplicate records, and unstructured inputs before they ever enter the pipeline, rather than after they have poisoned an output.

Consider a distributed retail organisation operating dozens of branches, in which the point-of-sale system, the customer relationship platform, and the finance ledger each maintain their own definition of a "customer" and their own version of "monthly revenue". For years this caused only mild friction, reports were reconciled by hand once a month, and the discrepancies were absorbed. The moment the organisation deployed a natural-language analytics agent, the friction became visible and acute. Asked a single question, "what was our revenue last quarter?", the agent could return three defensible but different answers, depending on which system it queried first. The failure was not in the model. It was in the absence of a governed schema and an authoritative source. Only once the organisation defined a canonical customer record and a single revenue definition, and aligned every downstream system to it, did the agent become trustworthy enough to inform a decision.

Section

Establishing Order in Files and File Management

Structured databases are only half the estate. Unstructured data, documents, PDFs, contracts, spreadsheets, and the sediment of internal communications, forms the backbone of most modern AI use cases, and it is invariably the messier half. If files are disorganised, an AI agent cannot retrieve or synthesise accurate context, however sophisticated its reasoning. Rigorous file management therefore becomes mandatory. This means taxonomy and naming conventions strict enough that an agent can parse content reliably from metadata alone; access and version control that guarantees the agent reads only the current, approved version of a document; and directory structures that map logically and hierarchically onto the organisation's actual operational workflows.

The stakes here are easily underestimated. Picture a firm that manages a large portfolio of tenancy and service agreements, with contracts accumulated over a decade across several shared drives and personal folders, named inconsistently and duplicated liberally. Deployed against this estate, a retrieval agent tasked with summarising a client's current obligations dutifully surfaced a superseded 2021 draft rather than the executed 2024 agreement, because nothing in the file's name, location, or metadata signalled which was authoritative. The agent behaved exactly as designed. It simply had no way to distinguish truth from residue. A single afternoon of disciplined version control and naming convention would have prevented an error that, in a compliance context, could have carried genuine consequence.

Section

Governance as Data Discipline

This restructuring is a monumental undertaking, and it cannot survive on goodwill or ad hoc guidelines. Serious programmes are building explicit AI governance structures, because the establishment of clean data pipelines touches every corner of the business. In practice this means cross-functional bodies with genuine representation from IT, data, security, legal, and the operating business, not a committee convened once to approve a policy, but a standing function that owns data quality as a continuous discipline.

A recurring finding from recent research is that successful AI implementations hinge on a fully engaged C-suite and board, reinforced by targeted workforce enablement. In the most mature organisations, leaders use AI themselves, not as a symbolic gesture, but because doing so signals importance and materially reduces resistance further down the organisation.

Here the argument meets its most persistent obstacle, which is not technical but human. Cultural resistance remains the primary impediment to disciplined data reform, rooted in a deep-seated preference for the familiar over the uncertain. Restructuring foundational information is unglamorous, laborious, and offers no immediate headline; it is far easier to sponsor a visible pilot than to fund the invisible plumbing on which every pilot depends. For personnel accustomed to grounding decisions on well-worn spreadsheets and trusted workarounds, submitting the organisation's information to a single governed standard represents a real intellectual and fiduciary adjustment.

Yet the sequence is not negotiable. Before an organisation can automate, predict, or converse with its own information, it must first organise, structure, and manage the data that fuels those ambitions. The enterprises that will extract genuine advantage from AI are not those that deploy it earliest or most aggressively. They are those with the discipline to sort their data first, and the patience to do the quiet work before reaching for the powerful tool.

YP

Yarnin Peled

Head of IT & Technology Projects | IMBA Candidate, Bar-Ilan University

Writing on digital transformation, operational excellence, and practical economics of AI.