Building Trustworthy Data Foundations for Artificial Intelligence
AI-ready data isn’t just about having more data—it’s about continuously ensuring the right data is diverse, timely, accurate, secure, discoverable, and usable within its specific context.
Sanchez P.
9/2/202693 min read


Abstract
The effectiveness, reliability and trustworthiness of artificial intelligence (AI) systems are fundamentally dependent on the quality, relevance and governance of the data on which they rely. As organisations increasingly adopt machine learning (ML), generative artificial intelligence (GenAI) and large language model (LLM) applications, conventional approaches to data management are insufficient to address the broader requirements of AI-enabled environments. AI systems require data that are not only accurate, but also appropriately diverse, timely, secure, discoverable and technically consumable. Building on Qlik's (2026) six principles of AI-ready data—diversity, timeliness, accuracy, security, discoverability and consumability—this paper develops an academic framework for understanding AI data readiness.
Drawing on the peer-reviewed literature on dataset quality, algorithmic bias, data governance and reproducibility, the paper critically examines each of the six principles and establishes their interdependence. The analysis demonstrates that diversity supports representativeness and can contribute to reducing certain forms of bias; timeliness supports relevance in changing environments; accuracy provides a foundation for reliable AI outputs; security supports privacy, integrity and responsible access; discoverability enables data to be identified, understood and governed; and consumability enables data to be appropriately transformed and incorporated into ML and GenAI architectures.
The paper argues that AI readiness should not be understood as a static property of a dataset or as a simple extension of traditional data quality. Instead, it should be conceptualised as a multidimensional, context-dependent and continuously maintained organisational capability. The proposed framework emphasises the interaction between the six dimensions and recommends that organisations assess them in relation to the intended AI application, operating environment and associated risks. It further distinguishes between minimum readiness thresholds and organisational maturity, recognising that critical weaknesses in areas such as security or accuracy should not necessarily be offset by strengths in other dimensions. The framework provides a basis for organisations to assess, remediate, monitor and continuously improve their data foundations for responsible and sustainable AI adoption.
Keywords: artificial intelligence; AI data readiness; data quality; data governance; trustworthy AI; machine learning; generative AI; data management.
1. Introduction
Artificial intelligence (AI) is increasingly embedded in organisational decision-making, scientific research and digital services. Machine learning (ML) systems are used to identify patterns, generate predictions and support classification and decision-making, while generative artificial intelligence (GenAI) and large language models (LLMs) are increasingly being applied to the production and transformation of text, software code, images and other forms of digital content. Despite these advances, the effectiveness and trustworthiness of AI systems remain fundamentally dependent on the quality, relevance and governance of the data on which they are developed and operated. AI should therefore not be understood solely as a problem of model selection or computational capability; it is equally a problem of establishing reliable and appropriately governed data foundations.
The importance of data quality to AI has been increasingly recognised in the literature. Priestley et al. (2023) demonstrate that dataset quality is a critical factor in machine learning and identify quality as a lifecycle concern extending from data collection and preparation through to model development and deployment. This perspective is significant because deficiencies introduced at earlier stages of the data lifecycle can propagate into subsequent stages of AI development and ultimately affect model performance. Similarly, Heil et al. (2021) emphasise that reproducible machine learning requires systematic attention to data collection, cleaning, curation and transformation, alongside appropriate documentation of models and computational processes. Data quality is therefore not a peripheral consideration but a fundamental condition for reliable and reproducible AI.
The relationship between data and AI is also a governance issue. Janssen et al. (2020) argue that effective data governance provides an important foundation for trustworthy AI by establishing structures through which data, processes and algorithms can be managed, scrutinised and held accountable. This is particularly important where AI systems influence decisions affecting individuals or organisations, because the consequences of poor data management may extend beyond technical model error to issues of fairness, privacy, accountability and public trust. The challenge is consequently not simply to make data available, but to ensure that data are appropriate for their intended AI application and governed throughout their lifecycle.
In this context, Qlik’s (2026) The Six Principles of AI-Ready Data provides a practical framework for considering the characteristics required for effective AI applications. The framework identifies six principles of AI-ready data: diversity, timeliness, accuracy, security, discoverability and consumability. Collectively, these principles challenge the assumption that the availability or volume of data is sufficient to support successful AI implementation. Instead, they suggest that data must possess a combination of representational, temporal, quality, governance and technical characteristics before they can provide a reliable foundation for AI.
The concept of AI data readiness therefore extends beyond conventional understandings of data quality. Traditional data quality concerns commonly emphasise characteristics such as accuracy, completeness, consistency and timeliness. These remain important, but AI systems introduce additional considerations concerning representativeness, provenance, documentation, accessibility, security and technical usability. Priestley et al. (2023) demonstrate that dataset quality encompasses multiple dimensions that can influence machine learning outcomes, while Janssen et al. (2020) highlight the importance of governance arrangements in establishing trustworthy AI. Together, these perspectives indicate that data suitability for AI cannot be reduced to the correctness of individual data values.
A further challenge arises from the fact that AI data requirements are highly dependent on context. Data that are adequate for descriptive reporting may be unsuitable for predictive modelling if they do not adequately represent the target population or relevant conditions. Likewise, data that were appropriate when a model was developed may become less relevant as the environment in which the model operates changes. This dynamic character of AI data is reinforced by the emphasis on reproducibility and lifecycle management in the literature (Heil et al., 2021). AI data readiness should consequently be understood as a contextual and evolving property rather than a permanent characteristic of a dataset.
The issue of diversity illustrates this distinction particularly clearly. AI systems can reflect patterns and assumptions embedded within their training data, meaning that limited or systematically unrepresentative datasets may contribute to biased outcomes. Bernhardt, Jones and Glocker (2022) demonstrate the complexity of identifying sources of dataset bias and caution against simplistic explanations of disparities in machine learning outcomes. Diversity should therefore not be interpreted merely as the accumulation of larger volumes of data or the inclusion of multiple data sources. Rather, it concerns whether the available data adequately represent the populations, conditions and circumstances relevant to the intended application.
At the same time, data must remain sufficiently current to support AI systems operating in changing environments. Changes in user behaviour, organisational processes, scientific knowledge and external conditions can reduce the relevance of previously collected data. Timeliness is therefore closely connected to the lifecycle perspective identified by Priestley et al. (2023) and Heil et al. (2021). The question is not simply whether data are recent, but whether their temporal characteristics remain appropriate for the decision or task for which an AI system is being used.
Security, discoverability and consumability introduce further dimensions of AI readiness. Security is necessary to protect sensitive information and maintain appropriate control over its use, while effective data governance supports responsible access, transparency and accountability (Janssen et al., 2020). Discoverability is important because data that cannot be identified, interpreted or appropriately documented are difficult to reuse and govern. This aligns with the broader emphasis on documentation and reproducibility in machine learning research (Heil et al., 2021; Nature Computational Science, 2021). Consumability, meanwhile, concerns whether data can be transformed and incorporated effectively into the technical architectures used by ML and GenAI systems. Thus, data availability alone does not establish data usability.
These considerations support the central argument of this paper: AI readiness should be conceptualised as a multidimensional and dynamic capability that integrates data quality, representativeness, temporal relevance, security, governance, discoverability and technical usability. The six principles proposed by Qlik (2026) provide a useful organising framework, but their significance is strengthened when considered alongside established academic research on dataset quality, algorithmic bias, data governance and reproducibility. The principles should therefore be viewed not as independent characteristics but as interconnected dimensions whose combined performance determines the suitability of data for AI applications.
Accordingly, this paper critically examines the six principles of AI-ready data and develops them into an academic framework for understanding AI data readiness. It first considers the relationship between conventional data quality and the broader requirements of AI. It then examines diversity and algorithmic bias, timeliness and changing AI environments, accuracy and reliability, security and data governance, discoverability and documentation, and consumability for ML and GenAI applications. Finally, the paper considers how these dimensions interact and why AI readiness should be assessed continuously throughout the AI lifecycle rather than treated as a one-time data quality assessment.
The overall proposition is that trustworthy AI depends not simply on having more data, but on having data that are fit for purpose, appropriately governed, technically usable and continuously evaluated. AI data readiness should therefore be understood as an organisational capability that connects the technical foundations of data management with the broader requirements of reliable, reproducible and trustworthy AI.
2. From Data Quality to AI Data Readiness
Traditional approaches to data quality provide an important foundation for understanding the requirements of artificial intelligence (AI), but they are insufficient to capture the full range of conditions necessary for effective and trustworthy AI applications. Conventional data quality frameworks typically emphasise dimensions such as accuracy, completeness, consistency and timeliness. These dimensions remain fundamental because errors, omissions and inconsistencies can directly affect the reliability of analytical outputs. However, the distinctive characteristics of machine learning (ML), generative artificial intelligence (GenAI) and large language models (LLMs) introduce additional requirements relating to representativeness, provenance, documentation, governance and technical usability. AI data readiness should therefore be understood as a broader concept than conventional data quality.
The distinction is important because AI systems do not simply retrieve or report information contained within a dataset. Machine learning systems identify statistical relationships and patterns from data and use these relationships to generate predictions, classifications or other outputs. Consequently, characteristics of the data that may have limited significance in conventional reporting environments can have substantial effects on AI behaviour. Priestley et al. (2023) demonstrate that dataset quality is a multidimensional issue in machine learning and that quality concerns can arise throughout the dataset lifecycle. Their analysis reinforces the view that data quality must be considered in relation to how data are collected, prepared, transformed and ultimately used by an AI system.
This lifecycle perspective is also evident in research on reproducibility. Heil et al. (2021) argue that reproducible machine learning requires attention not only to algorithms and models but also to the processes through which data are collected, cleaned and curated. Data transformations are particularly important because they can alter the characteristics and meaning of the original information. Without appropriate documentation of these processes, it becomes difficult to determine how particular datasets contributed to model outcomes or to reproduce the conditions under which a model was developed. Data readiness must therefore incorporate not only the quality of the data themselves but also the processes and information required to understand and reproduce their use.
A further distinction concerns the relationship between data quality and context. Data cannot be considered universally fit or unfit for AI; their suitability depends on the purpose, population, environment and model for which they are intended. A dataset may be accurate and internally consistent while nevertheless being inappropriate for a particular predictive task because it does not adequately represent the target population or relevant operating conditions. Similarly, historical data may accurately describe the environment in which they were collected but become less suitable when patterns of behaviour or external circumstances change. Data quality should therefore be assessed in relation to intended use rather than treated as an absolute property.
This contextual perspective is particularly important when considering bias. The presence of accurate observations does not necessarily mean that a dataset will support fair or generalisable AI outcomes. Dataset bias can arise from sampling, measurement processes, missing information and historical conditions, among other factors. Bernhardt, Jones and Glocker (2022) demonstrate that identifying the sources of bias in datasets can be complex and that observed differences in algorithmic outcomes cannot necessarily be attributed to a single feature of the data. This suggests that AI readiness must consider not only whether data are technically correct, but also whose experiences and characteristics are represented, how the data were generated and what limitations may affect their use.
The concept of provenance is consequently central to AI data readiness. Knowing where data originated, how they were collected and what transformations they underwent provides essential context for evaluating their suitability. Provenance also supports accountability when AI systems produce unexpected or problematic results. If an organisation cannot establish which data were used, how those data were transformed or which version of a dataset informed a model, it becomes difficult to investigate errors or reproduce outcomes. The emphasis on documentation and reproducibility in machine learning research therefore provides an important academic basis for treating provenance as a component of data readiness (Heil et al., 2021; Nature Computational Science, 2021).
AI also creates requirements that extend beyond the informational characteristics of a dataset to the organisational structures governing its use. Janssen et al. (2020) identify data governance as a foundational component of trustworthy AI and emphasise the need to organise data in ways that support responsible use, transparency and accountability. From this perspective, data readiness cannot be achieved solely through technical data-cleaning activities. Organisations must also establish appropriate responsibilities, controls and governance processes around the data. Questions concerning ownership, access, accountability, appropriate use and transparency become part of the conditions under which data can be considered ready for AI.
The emergence of GenAI and LLM applications further reinforces this distinction. Unlike some conventional analytical applications, GenAI systems can operate across large collections of structured and unstructured information and may require data to be transformed into formats suitable for retrieval, processing or model interaction. Consequently, data must not only be accurate but also sufficiently structured, documented and technically accessible to support the intended architecture. This creates a distinction between data availability and data usability. Information may exist within an organisation but remain practically unavailable to an AI system because it cannot be located, interpreted, accessed or transformed appropriately.
The six principles proposed by Qlik (2026)—diversity, timeliness, accuracy, security, discoverability and consumability—provide a useful framework for capturing these broader requirements. Accuracy and timeliness connect directly with established concerns in data quality, while diversity addresses representativeness and potential bias. Security introduces the requirements of protection and controlled access, discoverability addresses the ability to locate and understand organisational data, and consumability addresses their technical suitability for AI applications. Taken together, the six principles provide a broader conceptualisation of data readiness than a conventional focus on data correctness alone.
The value of this framework lies particularly in the relationships between its dimensions. The principles should not be interpreted as independent criteria that can be satisfied sequentially. Instead, weaknesses in one dimension can undermine the value of others. For example, highly accurate data may still be unsuitable if they are outdated or unrepresentative. Data that are diverse and current may nevertheless create significant risks if they are inadequately secured. Similarly, technically consumable data may be difficult to interpret or govern if they lack sufficient metadata and provenance. AI readiness is therefore better understood as a system of interdependent properties than as a collection of isolated data quality measures.
This integrated perspective also helps distinguish data quality from AI data readiness. Data quality primarily concerns whether data possess characteristics that make them reliable and fit for a defined purpose. AI data readiness encompasses these concerns but extends them to include whether data can be responsibly and effectively incorporated into an AI lifecycle. In conceptual terms, conventional data quality can therefore be viewed as a necessary component of AI readiness, but not a sufficient condition for it.
The distinction can be expressed through three broad requirements. First, data must be fit for purpose, meaning that they are sufficiently accurate, relevant, representative and timely for the intended AI application. Second, data must be governable, meaning that their provenance, security, ownership, documentation and appropriate use can be established and controlled. Third, data must be technically usable, meaning that they can be discovered, accessed, transformed and consumed by the relevant AI architecture. These requirements correspond closely with the six principles proposed by Qlik (2026) while providing a broader conceptual interpretation grounded in the academic literature.
AI readiness should also be viewed as a lifecycle property rather than a one-time assessment. The suitability of data can change as datasets are updated, organisational processes evolve, models are retrained and deployment environments change. Priestley et al. (2023) emphasise the lifecycle nature of dataset quality, while Heil et al. (2021) highlight the importance of documenting data-related processes to support reproducibility. These perspectives suggest that an organisation cannot establish AI readiness simply by validating a dataset before initial model development. Instead, readiness requires continuing assessment as data and AI systems evolve.
This is particularly significant because AI systems may create feedback between data and operational environments. Once deployed, an AI system may influence decisions or behaviours that subsequently affect the data generated by the organisation. Changes in these data can then influence future model development and performance. AI readiness must therefore accommodate the possibility that the relationship between data, models and their operating environment will evolve over time.
On this basis, AI data readiness can be conceptualised as a multidimensional, contextual and dynamic capability. It incorporates conventional data quality while extending beyond it to address representativeness, temporal relevance, security, governance, discoverability and technical usability. The six principles proposed by Qlik (2026) provide a practical structure through which these dimensions can be examined, while the academic literature establishes their broader significance for reliable, reproducible and trustworthy AI (Priestley et al., 2023; Heil et al., 2021; Janssen et al., 2020).
The conceptual shift from data quality to AI data readiness is therefore not merely terminological. It reflects a change in the role of data within contemporary AI systems. Data are no longer simply inputs to analytical processes; they form part of the technical, organisational and governance infrastructure through which AI systems operate. Understanding AI readiness in these broader terms provides the foundation for examining the six principles individually and assessing how they collectively contribute to trustworthy AI.
3. Data Diversity and the Problem of Algorithmic Bias
Data diversity is a central principle of AI data readiness because the characteristics of training and evaluation data can substantially influence the behaviour, reliability and generalisability of artificial intelligence (AI) systems. Qlik (2026) identifies diversity as one of the six principles of AI-ready data, emphasising the importance of drawing on sufficiently broad sources and patterns to support AI applications. However, diversity should not be interpreted simply as the quantity of data collected or the number of datasets combined. From an academic perspective, the more important question is whether data adequately represent the populations, circumstances, behaviours and conditions relevant to the intended AI application. Diversity is therefore closely connected to representativeness, dataset bias and the validity of AI outputs.
Machine learning (ML) systems learn statistical relationships from the data on which they are trained. Consequently, patterns that are overrepresented, underrepresented or systematically absent from a dataset can influence the relationships learned by a model. This creates a fundamental distinction between data volume and data diversity. A very large dataset may still provide an incomplete representation of the environment in which an AI system will operate. Conversely, a smaller dataset may contain meaningful variation across the relevant populations and conditions if it has been appropriately designed and collected. AI readiness should therefore assess whether the diversity contained within a dataset is relevant to the intended use rather than assuming that larger datasets are inherently superior.
The relationship between diversity and algorithmic bias is particularly important. Bias can enter datasets through multiple stages of the data lifecycle, including sampling, measurement, labelling, missing data and the historical circumstances under which information was generated. Bernhardt, Jones and Glocker (2022) demonstrate the complexity of identifying potential sources of dataset bias in medical AI and caution against attributing disparities in algorithmic outcomes to a single source. Their analysis illustrates that apparent bias may result from interactions between data collection processes, population characteristics and the conditions under which the data are generated. This complexity makes diversity an important but non-sufficient condition for reducing algorithmic bias.
A critical implication is that diversity does not automatically produce fairness. Increasing the number of demographic groups, data sources or observations represented within a dataset may improve coverage, but it does not necessarily remove structural or measurement biases. If the data generation process systematically reflects historical inequalities, inappropriate measurement practices or unequal access to services, simply increasing dataset diversity may reproduce those underlying patterns. Bernhardt, Jones and Glocker (2022) consequently provide an important basis for treating dataset bias as a problem requiring contextual investigation rather than as a simple deficiency that can be solved by adding more observations.
The concept of representativeness is therefore central to the assessment of AI-ready data. A dataset should be evaluated against the population and operating conditions for which the AI system is intended. This requires consideration of whether relevant demographic, geographic, behavioural, operational and contextual variation is adequately captured. Importantly, the appropriate dimensions of diversity will differ between applications. A dataset designed for a medical AI system may require different forms of representation from one designed for financial forecasting, customer analytics or organisational decision support. Diversity is consequently a context-dependent property, rather than a universal statistical threshold.
This context dependence also creates challenges when datasets are constructed from multiple sources. Combining datasets can increase the breadth of information available to an AI system, but it can also introduce inconsistencies in definitions, measurement practices and population coverage. Data originating from different environments may not be directly comparable, and apparent diversity may conceal differences in how variables were defined or recorded. Priestley et al. (2023) emphasise the multidimensional nature of dataset quality in machine learning, supporting the view that the addition of data sources should be accompanied by systematic assessment of their quality and suitability. Diversity should therefore be evaluated alongside accuracy, consistency, provenance and contextual relevance.
The issue of underrepresentation is particularly significant. Groups or conditions that occur less frequently within a dataset may contribute relatively little to the patterns learned by a model, even where they are operationally important. This can create a mismatch between statistical frequency and practical significance. A minority case may represent a small proportion of observations while carrying substantial consequences when an AI system makes an incorrect prediction. AI-ready data should therefore not be evaluated solely according to aggregate distributions; organisations should also identify whether critical populations, edge cases and operational scenarios are sufficiently represented.
At the same time, attempts to improve diversity through artificial balancing or oversampling require careful consideration. Modifying the distribution of a dataset may improve representation for a particular modelling task, but it can also change the relationship between variables or introduce additional assumptions. Consequently, interventions intended to improve diversity should themselves be documented and evaluated. The reproducibility literature emphasises the importance of documenting data collection, cleaning and curation processes so that the development of machine learning systems can be understood and reproduced (Heil et al., 2021). This principle applies directly to any intervention undertaken to address representational limitations.
Documentation is therefore an important complement to diversity. An AI-ready dataset should not only contain relevant variation but also provide information about how that variation was obtained and what limitations remain. Without documentation of sampling methods, collection conditions and known gaps, users may incorrectly assume that a dataset is representative when important groups or contexts are absent. Documentation can consequently help distinguish between known limitations and unknown limitations, improving the basis on which AI systems are evaluated and governed (Heil et al., 2021).
Diversity should also be considered across the AI lifecycle rather than solely during initial model development. A dataset may adequately represent the population at the time of development but become less representative as the deployment environment changes. Changes in user behaviour, population composition, organisational processes or external conditions can alter the relationship between the training data and the environment in which the model operates. This reinforces the broader argument developed in Chapter 2 that AI readiness is dynamic rather than a permanent attribute of a dataset.
The relationship between diversity and generalisability is also important. An AI system trained on narrowly defined conditions may perform effectively within those conditions but fail when exposed to previously underrepresented circumstances. A more diverse dataset can potentially improve the range of conditions captured by a model and therefore support greater robustness. However, diversity alone cannot guarantee generalisability. Model architecture, data quality, task definition and deployment conditions also influence performance. The appropriate conclusion is therefore that diversity is an important enabling condition for generalisability rather than a guarantee of it.
The principle can consequently be translated into a set of practical assessment requirements for AI-ready data. Organisations should consider whether datasets provide:
demographic representation, where demographic characteristics are relevant to the intended application;
contextual representation, including relevant environments, circumstances and operating conditions;
temporal variation, ensuring that data capture meaningful changes over time;
coverage of underrepresented groups, particularly where errors may have significant consequences;
appropriate sampling practices, with potential sources of selection and measurement bias examined;
documented dataset limitations, including known gaps and constraints; and
ongoing monitoring, to determine whether the representativeness of data changes after deployment.
These requirements demonstrate why diversity should be regarded as both a data quality and governance concern. From a quality perspective, insufficient representation can limit the reliability and generalisability of AI outputs. From a governance perspective, organisations need to understand whose data are represented, whose are missing and what consequences may result from those gaps. Janssen et al. (2020) argue that trustworthy AI requires governance arrangements that support transparency and accountability. Diversity therefore forms part of the broader governance question of whether an organisation can justify the data foundations of its AI systems.
There is also an important distinction between source diversity and population diversity. Drawing data from many sources does not necessarily mean that the resulting dataset represents the intended population. Multiple sources may contain similar biases, overlapping populations or comparable measurement limitations. Conversely, a small number of carefully selected sources may provide substantial coverage of relevant variation. AI readiness should therefore prioritise meaningful representational diversity over the superficial appearance of dataset heterogeneity.
The six-principle framework proposed by Qlik (2026) is useful in this context because it positions diversity alongside timeliness, accuracy, security, discoverability and consumability. This positioning is important: diversity should not be assessed independently from the other dimensions. For example, diverse data that are inaccurate may provide misleading representations; diverse data that are outdated may fail to reflect current populations; and diverse data that lack discoverability or provenance may be difficult to evaluate. The principle therefore derives much of its value from its interaction with the other dimensions of AI readiness.
A further implication is that organisations should avoid treating diversity as a single numerical target. A universal threshold for what constitutes a sufficiently diverse dataset would be difficult to justify because the relevant requirements depend on the AI task, target population and deployment environment. Instead, diversity assessment should be purpose-specific and evidence-based. Organisations should identify the populations and conditions that matter to the intended application, assess their representation within the data and document significant gaps and potential consequences.
The academic literature therefore supports a more nuanced interpretation of the diversity principle than the simple proposition that AI systems require more varied data. Dataset diversity is valuable because it can improve representation of relevant populations and conditions, potentially supporting robustness and reducing certain forms of systematic exclusion. However, diversity cannot by itself eliminate bias or guarantee fairness. The processes through which data are generated, measured, selected and transformed must also be examined (Bernhardt, Jones and Glocker, 2022; Priestley et al., 2023).
Consequently, data diversity should be understood as contextual representativeness supported by transparent data practices. An AI-ready dataset should provide sufficient coverage of the populations and circumstances relevant to its intended use, while its sources, limitations and construction processes should be documented and subject to ongoing review. This interpretation moves the diversity principle beyond a quantitative concern with data volume and establishes it as a foundational component of responsible and trustworthy AI.
Ultimately, the objective is not to maximise diversity in the abstract, but to ensure that the data provide an appropriate representation of the environment in which the AI system will operate. Such an approach recognises that algorithmic bias is often a consequence of complex interactions between data, context and system design rather than a simple property of individual variables. Treating diversity as a multidimensional and continuously assessed component of AI readiness therefore provides a stronger basis for developing AI systems that are reliable, appropriately generalisable and more capable of supporting responsible decision-making.
4. Data Timeliness and Dynamic AI Environments
Timeliness is a fundamental dimension of AI data readiness because the value of data is often dependent on their relationship to the time at which an artificial intelligence (AI) system is developed or deployed. Qlik (2026) identifies timeliness as one of the six principles of AI-ready data and emphasises the importance of maintaining current information through mechanisms such as low-latency data pipelines, change data capture and continuous data processing. While these mechanisms are technically important, timeliness should not be reduced to the speed at which data are collected or transmitted. For AI applications, the more important question is whether data remain temporally relevant to the environment, task and decision for which they are being used.
This distinction is particularly important because machine learning (ML) systems learn relationships from historical observations. A model may therefore continue to rely on patterns that were valid when its training data were collected but have subsequently changed. Priestley et al. (2023) identify dataset quality as a lifecycle issue in machine learning, supporting the need to consider the continuing suitability of data rather than assessing quality only at the point of initial preparation. Timeliness should consequently be understood as an ongoing property of the relationship between data and their intended application.
The challenge arises because different AI applications have different temporal requirements. In some contexts, information may remain useful for months or years, whereas in others its value may decline within minutes or even seconds. For example, a historical dataset may remain appropriate for identifying long-term patterns but be unsuitable for a decision requiring a current representation of rapidly changing conditions. Consequently, there is no universal definition of "timely" data. Timeliness must be evaluated in relation to the rate of change within the underlying environment and the decision horizon of the AI application.
This leads to an important distinction between data freshness and temporal relevance. Freshness concerns how recently information has been collected or updated. Temporal relevance concerns whether that information remains appropriate for the intended purpose. A recently updated dataset can still be unsuitable if it does not capture the variables or conditions relevant to the current decision. Conversely, historical information may remain highly valuable where the task requires understanding long-term trends or stable relationships. AI readiness should therefore avoid equating recency with quality.
Timeliness is also closely connected to the problem of data drift and changing environments. When the statistical characteristics of the environment represented in incoming data change over time, relationships learned from historical data may become less reliable. Even where the original dataset was accurate and representative, changes in behaviour, processes or external conditions can reduce its suitability for continued use. The lifecycle perspective identified by Priestley et al. (2023) therefore has particular significance for timeliness: data quality must be monitored as the operational context evolves rather than treated as a fixed attribute.
The reproducibility literature reinforces the importance of understanding when and how data were collected and transformed. Heil et al. (2021) emphasise the need to document data collection, cleaning and curation processes in machine learning research. Such documentation provides essential temporal context because it allows users to establish which version of a dataset was used, when it was collected and how it was subsequently transformed. Without this information, it becomes difficult to determine whether an apparent change in model performance is associated with the model itself, changes in the underlying data or changes in the environment.
Timeliness should therefore be considered across several dimensions.
4.1 Data Freshness
Data freshness refers to the time elapsed between the generation or collection of information and its availability for AI use. In applications where the operating environment changes rapidly, long delays between data generation and model consumption may reduce the usefulness of the information. Qlik (2026) consequently highlights mechanisms such as continuous processing and low-latency data pipelines as practical approaches to maintaining timely data.
However, minimising latency should not become an objective in isolation. Faster data processing does not necessarily produce better AI outcomes if the incoming information is incomplete, poorly validated or insufficiently contextualised. Timeliness must therefore remain connected to accuracy and other dimensions of AI readiness.
4.2 Temporal Relevance
Temporal relevance concerns whether the period represented by a dataset remains appropriate for the intended AI application. Historical information can be valuable when it captures persistent patterns, but its usefulness may decline when the underlying environment changes substantially. Organisations should therefore consider whether the age and temporal distribution of the data remain consistent with the decision problem.
This dimension also highlights the importance of distinguishing between historical validity and current applicability. Data can accurately represent past conditions without providing an accurate representation of present conditions. Consequently, an AI system may operate on technically accurate data while nevertheless producing outputs that are poorly aligned with the current environment.
4.3 Update Frequency
The appropriate frequency for updating data depends on the rate at which the underlying information changes and the consequences of using outdated information. Some applications may require continuous or near-real-time updates, while others may only require periodic refreshes. AI readiness should therefore establish an update frequency that reflects the temporal characteristics of the task rather than adopting a uniform organisational standard.
Update frequency should also be aligned with model retraining and evaluation processes. Updating an underlying dataset does not necessarily mean that an AI model immediately incorporates the new information. Organisations must therefore distinguish between the freshness of operational data and the frequency with which models or AI applications are updated.
4.4 Temporal Coverage
Timeliness also involves the extent to which a dataset captures meaningful variation across time. A dataset containing only recent observations may be highly current but insufficient for understanding seasonal patterns, longer-term trends or unusual conditions. Conversely, a dataset spanning a long historical period may provide valuable temporal coverage but contain information that is no longer relevant to the current operating environment.
Priestley et al. (2023) emphasise the multidimensional nature of dataset quality, supporting the view that temporal adequacy cannot be assessed through a single measure of recency. AI-ready data should instead provide an appropriate balance between historical coverage and current relevance.
4.5 Temporal Monitoring
Because AI environments can change after deployment, organisations should continuously monitor the relationship between incoming data and the conditions represented during model development. This is consistent with the broader lifecycle perspective of AI data quality and reproducibility (Priestley et al., 2023; Heil et al., 2021).
Monitoring should identify changes that could affect the continued suitability of data, including significant shifts in distributions, changes in operational processes and emerging patterns that were not adequately represented in historical data. The objective is not necessarily to eliminate all change but to identify when change becomes sufficiently significant to require investigation, data refresh or model reassessment.
4.6 Timeliness and Reproducibility
Timeliness introduces an important tension between adaptation and reproducibility. Continuously updating datasets may improve the relevance of an operational AI system, but it can also make it more difficult to reproduce the exact data environment associated with a previous model output. Heil et al. (2021) emphasise the importance of documenting data and computational processes to support reproducible machine learning. This suggests that organisations should maintain appropriate records of dataset versions, collection periods and transformations even when operational systems continuously ingest new information.
Reproducibility is therefore not necessarily opposed to continuously updated data. Rather, it requires organisations to preserve sufficient historical information to reconstruct the data conditions associated with particular models or decisions. This creates a governance requirement for versioning, provenance and documentation alongside mechanisms for data refresh.
4.7 Timeliness in Generative AI and Large Language Model Applications
Timeliness has particular significance for generative AI (GenAI) and large language model (LLM) applications. These systems may generate outputs based on information that is no longer current or may lack access to organisation-specific knowledge that has changed since model development. Qlik (2026) identifies retrieval-augmented generation (RAG) as an approach through which LLM applications can incorporate external or organisational information into their responses.
In a RAG-based architecture, the timeliness of the underlying knowledge base can be as important as the capabilities of the LLM itself. If documents, policies, procedures or organisational information are outdated, retrieval may provide information that is technically relevant to the query but operationally incorrect. Conversely, regularly updating the knowledge base can improve the currency of retrieved information without necessarily requiring the underlying language model to be retrained.
This distinction reinforces the broader argument that AI data readiness concerns the entire data and AI architecture rather than the model alone. For GenAI applications, the currency of documents, metadata, indexes and retrieval systems may all influence the relevance of generated outputs.
4.8 The Limits of Timeliness as an Independent Principle
Although timeliness is essential, it should not be treated as an independent indicator of AI readiness. More recent data are not inherently better than older data. A newly collected dataset may contain measurement errors, insufficient representation or incomplete information, while a historical dataset may provide high-quality evidence for a particular analytical purpose. Timeliness must therefore be evaluated in combination with accuracy, diversity, discoverability and consumability.
For example, a highly accurate dataset may lose operational value if it is no longer representative of current conditions. Conversely, a highly current dataset may produce unreliable AI outputs if it has not been adequately validated. Similarly, timely data that cannot be discovered, interpreted or technically consumed cannot provide an effective foundation for AI. The six principles proposed by Qlik (2026) should therefore be understood as interdependent dimensions rather than isolated measures.
4.9 Implications for AI Data Readiness
The analysis suggests that organisations should assess timeliness through a combination of technical, contextual and governance measures. An AI-ready data environment should establish:
the expected freshness of data for each AI application;
the acceptable delay between data generation and use;
the temporal relevance of historical observations;
appropriate data refresh and update frequencies;
sufficient temporal coverage for the intended task;
mechanisms for detecting significant changes in the data environment;
versioning and provenance to support reproducibility; and
processes for reassessing datasets when operational conditions change.
These requirements demonstrate that timeliness is fundamentally a contextual and lifecycle-based property. The appropriate standard depends on how rapidly the relevant environment changes, how frequently decisions are made and what consequences may arise from using outdated information.
The principle also has direct implications for organisational governance. Janssen et al. (2020) emphasise the importance of data governance in establishing trustworthy AI. Governance arrangements should therefore specify who is responsible for determining whether data remain sufficiently current, how changes are monitored and when datasets or AI systems should be reassessed. Timeliness should not be left solely to technical infrastructure; it requires organisational ownership and clearly defined expectations.
Ultimately, data timeliness should be understood not as a race towards real-time processing but as the alignment between the temporal characteristics of data and the temporal requirements of an AI application. This interpretation provides a more rigorous basis for assessing AI readiness because it recognises that freshness, relevance, temporal coverage and monitoring address different aspects of the problem.
The central implication is that AI-ready data must remain appropriately aligned with the environment in which an AI system operates. Because environments change, this alignment cannot be guaranteed permanently. Timeliness must therefore be continuously evaluated alongside accuracy, diversity, security, discoverability and consumability. In this sense, timeliness provides an important illustration of the broader proposition advanced in this paper: AI data readiness is not a one-time property of a dataset, but an ongoing capability to maintain the suitability of data as AI systems and their operating environments evolve.
5. Accuracy, Data Quality and Reliability
Accuracy is a fundamental principle of AI data readiness because artificial intelligence (AI) systems depend on data to learn patterns, estimate relationships and generate outputs. Qlik (2026) identifies accuracy as one of the six principles of AI-ready data and emphasises practices such as data profiling, remediation, continuous monitoring, data lineage and impact analysis. These practices reflect an important premise: the reliability of an AI system is constrained by the quality of the information on which it depends. However, accuracy should not be interpreted narrowly as the absence of incorrect individual values. For AI applications, data quality is multidimensional and must be assessed in relation to the intended task, context and lifecycle of the system.
Priestley et al. (2023) identify dataset quality as a major factor affecting machine learning (ML) systems and demonstrate that quality encompasses multiple dimensions rather than a single measure of correctness. This is particularly significant for AI because models can learn systematic patterns from data even when individual observations appear plausible. Errors, omissions, inconsistencies or inappropriate representations can therefore become embedded within the relationships learned by a model. Data quality should consequently be considered before and throughout model development rather than treated as a final validation step.
The conventional definition of accuracy generally concerns whether a recorded value correctly represents the phenomenon it is intended to describe. Although this remains essential, AI applications require a broader interpretation. A dataset may contain individually accurate observations while still being unsuitable for a particular AI task because it is incomplete, outdated, inconsistently structured or unrepresentative. Accuracy should therefore be distinguished from fitness for purpose. The relevant question is not simply whether the data are correct, but whether they are sufficiently reliable and appropriate for the intended AI application.
This distinction is important because AI systems can transform relatively small data deficiencies into systematic model behaviour. If a particular variable is consistently measured in a way that differs across groups or contexts, a model may interpret the resulting differences as meaningful patterns. Similarly, missing observations may not be randomly distributed, creating a distorted representation of the underlying population. Bernhardt, Jones and Glocker (2022) demonstrate the complexity of dataset bias and highlight the need to consider how data-generation processes can influence apparent disparities in AI outcomes. Accuracy must therefore be assessed alongside representativeness and the processes through which data are generated.
5.1 Accuracy as a Multidimensional Concept
For the purposes of AI data readiness, accuracy should be considered alongside several related dimensions of data quality. These include:
correctness – whether individual values accurately represent the intended information;
completeness – whether sufficient information is available for the intended task;
consistency – whether equivalent information is represented coherently across datasets and systems;
validity – whether values conform to defined formats, rules or permissible ranges;
relevance – whether the information contributes appropriately to the intended AI task; and
reliability – whether the data can be trusted as a sufficiently stable basis for AI use.
These dimensions are interconnected rather than independent. A dataset may be highly accurate at the individual-record level while remaining incomplete for the intended prediction task. Similarly, data can be valid according to technical rules while being conceptually inappropriate for the model. Priestley et al. (2023) support this broader view by treating dataset quality as a multidimensional problem within the machine learning lifecycle.
Consequently, an assessment of accuracy should move beyond simple error detection. Organisations should evaluate whether the data provide an adequate representation of the information required by the AI system and whether known limitations could affect model behaviour.
5.2 Data Profiling and Validation
Data profiling provides an important mechanism for identifying potential quality problems before data are incorporated into AI systems. Profiling can reveal missing values, anomalous distributions, duplicate records, unexpected formats and inconsistencies between variables or datasets. Qlik (2026) identifies profiling and remediation as practical mechanisms for strengthening AI data quality.
Validation should then establish whether identified data conform to appropriate quality requirements. This may include technical validation rules, range checks, structural checks and consistency assessments. However, validation rules must themselves be appropriate to the intended use. A dataset can satisfy formal validation requirements while still containing information that is misleading or insufficient for the AI task.
This reinforces the importance of combining automated quality controls with contextual assessment. Domain knowledge may be required to determine whether a particular value is plausible, whether a variable adequately captures the intended concept or whether a missing observation has substantive significance. AI data readiness therefore requires interaction between technical data-quality processes and knowledge of the domain in which the data are used.
5.3 Completeness and Missing Data
Completeness is particularly important because missing information can influence the patterns learned by ML systems. The absence of data is not necessarily neutral. Where missingness is associated with particular populations, conditions or operational processes, simply removing incomplete observations may change the composition of the dataset.
This issue demonstrates why completeness cannot be evaluated solely through the proportion of missing values. The significance of missingness depends on what is missing, why it is missing and how its absence affects the intended AI application. A small amount of missing information in a highly influential variable may be more consequential than a larger amount of missing information in a less important variable.
AI-ready data should therefore include mechanisms for identifying and documenting important missingness patterns. Where remediation or imputation is undertaken, the process should also be documented because transformations can affect the statistical properties and interpretation of the resulting dataset. The importance of documenting data preparation and curation processes is consistent with the reproducibility requirements identified by Heil et al. (2021).
5.4 Consistency and Data Integration
Consistency becomes particularly important when AI applications draw information from multiple organisational systems. Differences in naming conventions, definitions, formats or measurement procedures can create inconsistencies that are not immediately apparent at the individual-record level.
For example, equivalent concepts may be represented differently across systems, or the same variable may have different meanings depending on the source from which it originates. When such information is integrated into an AI pipeline, the resulting dataset may appear complete and technically valid while containing semantic inconsistencies.
Priestley et al. (2023) emphasise the multidimensional nature of dataset quality, supporting the need to consider consistency alongside other quality dimensions. Organisations should therefore establish appropriate definitions and validation mechanisms when combining data sources. Data integration should not be assumed to improve quality simply because it increases the amount of available information.
5.5 Data Leakage and Inappropriate Information
Accuracy must also be considered in relation to the timing and appropriateness of information provided to an AI model. Data that are factually correct can nevertheless produce misleading model performance if they contain information that would not legitimately be available at the time a prediction or decision is made.
This illustrates a broader principle of AI data readiness: technically correct data are not necessarily operationally appropriate data. The inclusion of information that is unavailable in the intended deployment context can create an unrealistic representation of model performance. Data preparation must therefore consider not only whether information is correct but whether its use is appropriate within the intended decision environment.
This requirement connects accuracy with the principle of timeliness discussed in Chapter 4. Data must be sufficiently current and temporally appropriate, while also ensuring that the information available to the model reflects the conditions under which the system will actually operate. Accuracy should therefore be assessed within the temporal and operational context of the AI application.
5.6 Data Lineage and Provenance
Data lineage is an important component of reliable AI because it enables organisations to trace information from its original source through subsequent transformations and into the systems that consume it. Qlik (2026) identifies lineage as a mechanism for improving confidence in data quality and supporting impact analysis.
The importance of lineage extends beyond technical troubleshooting. If an AI system generates an unexpected or problematic output, organisations need to establish which data contributed to the result and how those data were processed. Without adequate lineage, it can be difficult to identify whether the problem originated in the source data, a transformation process, an integration step or the AI model itself.
This requirement is closely connected to reproducibility. Heil et al. (2021) emphasise that reproducible machine learning requires appropriate documentation of data collection, cleaning and curation processes. Similarly, Nature Computational Science (2021) highlights reproducibility as an important concern in machine learning. Maintaining lineage and provenance therefore provides not only operational value but also an important basis for accountability and independent examination.
5.7 Continuous Monitoring and Data Quality
Data quality should not be assessed only before an AI system is deployed. Changes in source systems, business processes, user behaviour or external conditions can introduce new errors or alter the characteristics of a dataset. A dataset that satisfies quality requirements at one point in time may therefore require reassessment later.
Continuous monitoring provides a mechanism for identifying such changes. Monitoring can be used to detect anomalies, changes in distributions, unexpected missingness and other indicators that may signal deterioration in data quality. This lifecycle approach is consistent with the findings of Priestley et al. (2023), who position dataset quality as an ongoing concern within machine learning.
Monitoring should nevertheless be proportionate to the intended use and risk of the AI application. Not every variable requires the same monitoring frequency or level of scrutiny. Critical data elements that directly influence important decisions may require more intensive controls than information with limited influence on system outcomes.
5.8 Accuracy and Model Reliability
The relationship between data accuracy and model reliability is important but should not be overstated. High-quality data can provide a stronger foundation for AI, but accurate data do not guarantee accurate model outputs. Model architecture, task definition, training procedures and deployment conditions can also influence performance.
Conversely, poor data quality can constrain model performance even when the underlying model is technically sophisticated. This creates an important asymmetry: improvements to model architecture cannot necessarily compensate for fundamental deficiencies in the data used to train or operate the system.
The appropriate conclusion is therefore that data quality is a necessary foundation for reliable AI but not a sufficient guarantee of reliability. This distinction strengthens the broader AI readiness framework by positioning accuracy as one essential dimension within an interconnected system rather than as a single determinant of AI performance.
5.9 Accuracy and Algorithmic Bias
Accuracy also interacts with diversity and bias. A dataset can be factually accurate while systematically underrepresenting particular groups or contexts. In such circumstances, improving the correctness of individual observations will not necessarily address the underlying representational problem.
Bernhardt, Jones and Glocker (2022) demonstrate why dataset bias requires careful investigation of the processes through which data are generated and represented. This suggests that accuracy assessments should consider whether data are not only correct but also appropriately representative of the conditions in which the AI system will be used.
The implication is that accuracy and diversity are complementary rather than interchangeable. Accuracy concerns whether information correctly reflects what it purports to represent, whereas diversity concerns whether the dataset adequately captures the relevant range of populations and conditions. Both are required for a robust AI data foundation.
5.10 Accuracy, Governance and Accountability
Data quality also has an organisational dimension. Janssen et al. (2020) argue that data governance is an important foundation for trustworthy AI because organisations require structures that support responsible data management, transparency and accountability.
From this perspective, accuracy cannot be treated solely as the responsibility of data engineers or technical teams. Organisations need clear ownership of critical data, defined quality requirements and processes for responding to identified quality problems. Where data are used to support consequential AI applications, accountability should include an understanding of who is responsible for assessing data suitability and addressing material deficiencies.
This makes data governance an enabling condition for sustained data quality. Technical validation can identify problems, but governance determines how those problems are prioritised, addressed and communicated.
5.11 Assessing Accuracy within AI Data Readiness
The analysis suggests that an AI-ready approach to accuracy should incorporate several complementary activities:
profiling datasets before AI use;
defining quality requirements according to the intended application;
validating data against technical and contextual rules;
assessing completeness and patterns of missingness;
identifying duplication and inconsistency;
evaluating the appropriateness of integrated data sources;
documenting transformations and remediation processes;
maintaining data lineage and provenance;
monitoring quality throughout the AI lifecycle; and
reassessing quality when data sources, models or operating environments change.
These activities move the concept of accuracy beyond simple error correction towards a broader assessment of data fitness, reliability and traceability.
Importantly, these requirements should not be applied uniformly to every dataset. The appropriate level of quality assurance depends on the intended use, the consequences of error and the characteristics of the AI system. An experimental application may tolerate greater uncertainty than a system supporting high-impact organisational decisions. AI data readiness should therefore incorporate a degree of proportionality in which quality requirements reflect the significance and risk of the intended application.
5.12 Accuracy as an Interdependent Principle
The position of accuracy within the six-principle framework proposed by Qlik (2026) further demonstrates why no individual data characteristic can establish AI readiness. Accurate data that are outdated may produce irrelevant outputs; accurate data that lack diversity may produce systematically unrepresentative results; accurate data that are insecure may create privacy and governance risks; and accurate data that cannot be discovered or consumed effectively may remain practically unusable.
Accuracy should therefore be understood as a foundational but interdependent dimension of AI readiness. It provides the basis on which other data characteristics can deliver value, but it cannot compensate for deficiencies elsewhere in the data ecosystem.
The central argument of this chapter is consequently that accuracy should be interpreted as more than factual correctness. For AI systems, reliable data must be sufficiently correct, complete, consistent, valid, relevant and traceable for the intended application. These characteristics must also be maintained throughout the data and AI lifecycle.
Ultimately, AI-ready data require a shift from the question “Are these data accurate?” to the broader question “Are these data sufficiently reliable and fit for this AI purpose, and can their quality be demonstrated and maintained?” This shift is consistent with the multidimensional understanding of dataset quality proposed by Priestley et al. (2023), the emphasis on reproducibility and documentation identified by Heil et al. (2021), and the governance requirements highlighted by Janssen et al. (2020). Accuracy is therefore best understood not as a standalone technical property, but as a continuously governed component of the broader capability required to develop reliable and trustworthy AI.
6. Security, Privacy and Data Governance
Security is a fundamental dimension of AI data readiness because artificial intelligence (AI) systems increasingly depend on large and interconnected collections of organisational, personal and commercially sensitive information. Qlik (2026) identifies security as one of the six principles of AI-ready data and emphasises mechanisms such as data classification, protection and access control. These measures are important because inappropriate access to, disclosure of or manipulation of data can compromise not only confidentiality but also the reliability and trustworthiness of AI systems. However, security should not be understood solely as a technical cybersecurity function. In the context of AI, it is closely connected to privacy, data governance, accountability and the controlled use of information.
The significance of security is reinforced by the broader literature on trustworthy AI. Janssen et al. (2020) argue that data governance provides an important foundation for trustworthy artificial intelligence by establishing organisational structures through which data can be managed responsibly. From this perspective, securing AI data is not simply a matter of preventing unauthorised access. Organisations must also establish who is permitted to access particular data, for what purposes, under which conditions and with what level of accountability. Security therefore forms part of a wider governance framework concerned with responsible data use.
The relationship between security and AI is particularly important because AI applications can bring together information from multiple sources and make that information available through new forms of interaction. Machine learning (ML), generative artificial intelligence (GenAI) and large language model (LLM) applications may process structured databases, documents and other forms of organisational knowledge within integrated technical environments. This can increase the value of organisational data, but it can also increase the consequences of inappropriate access or inadequate controls.
Consequently, an AI-ready data environment should seek to establish a balance between protection and controlled accessibility. Data that are insufficiently protected may expose individuals or organisations to significant risks. At the same time, data that are excessively restricted may become inaccessible to legitimate users and AI systems, reducing their practical value. Security should therefore not be conceptualised as maximising restriction, but as establishing appropriate controls that enable authorised and accountable use.
6.1 Security as a Component of AI Data Readiness
Qlik (2026) places security alongside diversity, timeliness, accuracy, discoverability and consumability. This positioning is important because security interacts directly with the other dimensions of AI readiness. Accurate data are of limited organisational value if they cannot be accessed appropriately, while highly discoverable data may create significant risks if sensitive information can be located and accessed without adequate controls.
Security should therefore be evaluated according to the intended use of the data and the consequences associated with inappropriate access. The level of protection required for a dataset containing publicly available information may differ substantially from that required for personal, confidential or commercially sensitive information.
This suggests that AI data readiness requires a risk-based approach to security. Organisations should identify the sensitivity and importance of their data, assess the risks associated with their intended AI use and implement controls proportionate to those risks.
6.2 Data Classification
Data classification provides a foundation for applying appropriate security controls. Organisations need to understand what types of information they possess and distinguish between data according to factors such as sensitivity, confidentiality and potential consequences of misuse.
Classification is particularly important in AI environments because datasets that appear relatively harmless in isolation may become more sensitive when combined with information from other sources. The integration of datasets can increase the amount of information that can be inferred about individuals, organisational activities or business processes.
An AI-ready environment should therefore establish clear categories for data and corresponding requirements for access, storage, processing and sharing. Classification should also be maintained as datasets evolve, because changes in data content or combinations of information may alter their sensitivity.
This connects security directly with discoverability. Organisations cannot effectively control information that they cannot identify. Conversely, making data highly discoverable without understanding their sensitivity may increase the likelihood of inappropriate use. Security and discoverability should therefore be designed as complementary rather than opposing capabilities.
6.3 Access Control and Authorisation
Access control is another central requirement of secure AI data. Appropriate controls should ensure that users and systems can access only the information necessary for their legitimate activities. This principle is particularly important where AI applications aggregate information from multiple organisational systems.
The objective should not be to prevent access generally, but to establish authorised and accountable access. Users should have clearly defined permissions, while organisations should maintain mechanisms for monitoring how sensitive information is accessed and used.
Janssen et al. (2020) emphasise the organisational importance of data governance for trustworthy AI. Access controls therefore need to be supported by clear responsibilities and governance processes. Technical permissions alone are insufficient if organisations cannot determine who is accountable for granting access, reviewing permissions or responding to inappropriate use.
6.4 Privacy and the Responsible Use of Data
Privacy represents a particularly important dimension of security where AI systems process information relating to individuals. The ability of AI systems to combine and analyse large volumes of information can increase the potential value of data but may also increase the risks associated with inappropriate disclosure or inference.
Privacy should therefore be considered from the beginning of the data and AI lifecycle rather than introduced as a final compliance measure. Organisations should consider whether information is appropriate for the intended AI application, how it will be accessed and transformed, and whether unnecessary personal information can be excluded or protected.
The governance perspective developed by Janssen et al. (2020) is relevant here because responsible AI requires organisational arrangements that support appropriate data use and accountability. Privacy is therefore not simply a property of the database; it is also a consequence of how data are governed and incorporated into AI systems.
Where appropriate, organisations may use techniques such as masking or tokenisation to reduce exposure of sensitive information while retaining sufficient utility for legitimate AI applications. However, such transformations should themselves be documented and evaluated because they can affect the usefulness, interpretation or completeness of the resulting data.
6.5 Data Protection Across the AI Lifecycle
Security requirements extend across the entire AI data lifecycle. Protection is required when data are collected, stored, transferred, transformed, used for model development and incorporated into operational AI applications.
This lifecycle perspective is consistent with the broader argument developed in earlier chapters. Priestley et al. (2023) demonstrate that dataset quality is a lifecycle concern, while Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes for reproducible machine learning. The same principle applies to security: controls should not be limited to a single stage of the data lifecycle.
For example, information may be appropriately secured within a source database but become exposed when copied into an analytical environment or incorporated into an AI application. Security assessments should therefore follow the movement and transformation of data rather than focusing only on their original storage location.
6.6 Security and Data Integrity
Security also has an important relationship with data integrity. AI systems depend not only on the confidentiality of data but also on confidence that information has not been improperly modified or manipulated.
If data used to train or operate an AI system are altered without appropriate controls, the resulting model or output may no longer reflect the intended information. This means that security failures can become data-quality failures and, ultimately, AI reliability failures.
The distinction is therefore useful between:
confidentiality, which concerns preventing inappropriate disclosure;
integrity, which concerns maintaining the correctness and trustworthiness of information; and
availability, which concerns ensuring that authorised users and systems can access information when required.
All three dimensions contribute to AI data readiness. Protecting confidentiality while allowing information to become unavailable to legitimate AI applications would not constitute effective data readiness. Similarly, maintaining availability without protecting data integrity could undermine the reliability of AI outputs.
6.7 Security, Generative AI and Large Language Models
The emergence of GenAI and LLM applications creates additional considerations for data security because organisational information may be incorporated into systems that generate responses based on retrieved or processed information. Qlik (2026) identifies retrieval-augmented generation (RAG) as an approach for connecting LLM applications with organisational knowledge.
In such environments, security controls must extend beyond the underlying data repository to the mechanisms through which information is retrieved and presented to the AI system. Sensitive information may be exposed if access controls are not appropriately applied to retrieval processes or if users can obtain information through AI interfaces that they would not otherwise be authorised to access.
This illustrates a broader principle: security controls must follow the data into the AI architecture. It is insufficient to secure the original source if subsequent processing, indexing, transformation or retrieval creates new pathways through which sensitive information can be accessed.
The same principle applies to the use of organisational documents in GenAI systems. Before information is incorporated into a retrieval or knowledge environment, organisations should establish whether it is appropriate for that purpose and what access restrictions should remain in place.
6.8 Security and Discoverability: A Necessary Tension
Security and discoverability may initially appear to be competing principles. Discoverability encourages organisations to make data easier to locate and understand, whereas security may require restrictions on access. However, these principles can be reconciled through controlled discoverability.
Users may need to know that a dataset exists without necessarily being authorised to view its underlying content. Metadata can therefore be used to provide information about the existence, ownership, purpose and sensitivity of data while maintaining appropriate restrictions over the data themselves.
This distinction reinforces the argument developed in Chapter 7 that metadata and documentation are not merely administrative mechanisms. They can support both effective discovery and responsible governance by enabling organisations to understand their data landscape without necessarily exposing sensitive information.
6.9 Security and Data Governance
The relationship between security and governance is central to the concept of AI readiness. Janssen et al. (2020) position data governance as an organisational foundation for trustworthy AI, emphasising the need for structures through which data can be responsibly organised and managed.
From this perspective, security should be embedded within governance arrangements that establish:
ownership and responsibility for critical datasets;
data classification requirements;
access and authorisation policies;
procedures for reviewing permissions;
requirements for monitoring data access and usage;
processes for responding to security incidents;
controls governing data sharing and movement; and
accountability for the use of data within AI applications.
These arrangements transform security from an isolated technical function into an organisational capability. The objective is not simply to secure data, but to establish demonstrable and accountable controls over how data are used.
6.10 Governance, Accountability and Trustworthy AI
The governance dimension becomes particularly important when AI systems influence organisational decisions. Organisations need to be able to demonstrate that the data used by AI systems were appropriately controlled and that responsibilities for their use were clearly defined.
Janssen et al. (2020) argue that trustworthy AI requires appropriate governance arrangements. This supports the proposition that data security contributes to trust not merely because it reduces technical risk, but because it provides evidence that an organisation exercises responsible control over its information assets.
Documentation also supports this accountability. Heil et al. (2021) emphasise the importance of documenting data-related processes to support reproducibility. Although their primary focus is reproducibility in machine learning, the underlying principle is relevant to security governance: organisations should be able to establish how data were obtained, processed and used.
6.11 The Limits of Security as an Independent Principle
Security is essential, but like the other dimensions of AI readiness, it cannot be evaluated independently. Highly secure data may still be unsuitable for AI if they are inaccurate, outdated or unrepresentative. Conversely, highly accurate and diverse data may create unacceptable risks if sensitive information is inadequately protected.
There is also a potential trade-off between protection and usability. Excessive restrictions may prevent legitimate AI applications from accessing valuable information, while insufficient controls may expose sensitive data. AI readiness therefore requires an appropriate balance between security, accessibility and utility.
This balance is particularly important because the value of AI depends partly on the ability to use relevant organisational information effectively. Security should not therefore be interpreted as a requirement to minimise access wherever possible. Instead, it should support controlled access that is appropriate to the sensitivity and purpose of the data.
6.12 Assessing Security within AI Data Readiness
An AI-ready security framework should therefore consider:
classification of data according to sensitivity and risk;
clearly defined ownership and accountability;
appropriate authentication and authorisation;
access controls based on legitimate requirements;
protection of data during storage and transfer;
mechanisms for maintaining data integrity;
appropriate masking or tokenisation where required;
monitoring of access and data usage;
controls over data movement between systems;
security requirements for AI retrieval and processing environments; and
periodic reassessment as datasets, AI applications and organisational conditions change.
These measures should be proportionate to the intended use and potential consequences of misuse. Security requirements for experimental or low-risk applications may differ from those for AI systems processing sensitive organisational or personal information.
6.13 Security as an Interdependent Principle
The position of security within the six-principle framework proposed by Qlik (2026) reinforces the argument that AI readiness is multidimensional. Security cannot compensate for poor accuracy, nor can accurate data compensate for inadequate protection. Similarly, discoverability must be balanced with controlled access, while consumability must be achieved without creating inappropriate exposure of sensitive information.
Security is therefore best understood as an enabling governance condition for responsible data use. Its purpose is not simply to restrict access but to establish appropriate conditions under which data can be accessed, processed and consumed by authorised people and systems.
The central argument of this chapter is that AI-ready data must be not only accurate, diverse, timely and technically usable, but also appropriately protected and governed throughout their lifecycle. Security, privacy and governance should consequently be treated as interconnected rather than separate concerns.
Ultimately, trustworthy AI requires organisations to demonstrate that they can control both who can access data and how those data are used within AI systems. This moves the concept of security beyond conventional perimeter protection towards a broader model of accountable data governance. In this sense, security is not an obstacle to AI readiness; when appropriately designed, it is one of the conditions that makes responsible AI use possible.
7. Discoverability, Metadata and Organisational Knowledge
Discoverability is a critical but often underappreciated dimension of AI data readiness. Qlik (2026) identifies discoverability as one of the six principles of AI-ready data and emphasises the importance of metadata, semantic typing, business glossaries and searchable data catalogues. The underlying premise is that data cannot contribute effectively to artificial intelligence (AI) applications if organisations are unable to identify, locate, interpret and appropriately access them. In this respect, discoverability extends beyond the technical ability to find a dataset. It encompasses the contextual information required to understand what data represent, where they originated, how they were created and how they should appropriately be used.
The importance of discoverability becomes increasingly apparent as organisations accumulate large volumes of data across multiple systems, departments and platforms. Data may exist within an organisation without being visible to potential users, or may be technically accessible but poorly understood. In either situation, the practical value of the data is reduced. For AI applications, this problem can be particularly significant because machine learning (ML), generative artificial intelligence (GenAI) and large language model (LLM) systems may depend on the ability to identify relevant information across heterogeneous data environments.
Discoverability is therefore closely connected to the broader concept of data usability. A dataset that can technically be accessed but whose meaning, provenance or limitations are unclear cannot necessarily be considered AI-ready. Users need sufficient contextual information to determine whether the data are appropriate for a particular application. This connects discoverability directly with the principles of accuracy, diversity, timeliness and security discussed in earlier chapters.
7.1 From Data Availability to Data Discoverability
A fundamental distinction should be made between data availability and data discoverability. Data availability concerns whether information exists and can technically be accessed. Discoverability concerns whether users or systems can identify relevant information and understand its potential value and limitations.
This distinction is important because organisations may possess substantial quantities of information that remain effectively invisible to AI development teams. Data may be stored in departmental systems, legacy applications or isolated repositories without adequate cataloguing or documentation. As a result, teams may duplicate data collection, rely on incomplete sources or use datasets that are convenient rather than genuinely appropriate.
Qlik (2026) addresses this challenge through the use of searchable data catalogues, metadata and business glossaries. These mechanisms can provide a structured representation of an organisation's data landscape and enable users to identify potentially relevant information more efficiently.
Discoverability should therefore be regarded as an organisational capability rather than simply a feature of a data platform.
7.2 Metadata as the Foundation of Discoverability
Metadata provide the information required to describe and contextualise datasets. Technical metadata may describe structures, fields, formats and relationships, while business metadata may explain the meaning and organisational purpose of particular data elements.
For AI applications, metadata can provide critical information about:
the origin of a dataset;
its ownership and stewardship;
the meaning of variables and fields;
the period covered by the data;
update frequency;
data quality characteristics;
known limitations;
relationships with other datasets; and
appropriate or inappropriate uses.
Without such information, users may be able to locate a dataset but remain unable to determine whether it is suitable for their intended AI application.
Metadata therefore performs a function beyond administrative description. It provides the context necessary for evaluating data fitness for purpose.
7.3 Semantic Meaning and Interpretation
Discoverability also has a semantic dimension. AI applications increasingly operate across data that originate from different systems and organisational contexts. The same term may have different meanings across departments, while different terms may be used to describe equivalent concepts.
Business glossaries and semantic definitions can help address this problem by establishing shared meanings for important organisational concepts. Qlik (2026) identifies semantic typing and business glossaries as components of discoverability, highlighting the importance of enabling both people and systems to interpret data consistently.
This is particularly relevant to AI because models and retrieval systems require meaningful relationships between concepts. If the semantic meaning of a dataset is unclear, technically sophisticated AI systems may still produce inappropriate results because the underlying information has been misunderstood or incorrectly mapped to the intended task.
Discoverability should therefore include not only the ability to locate data but also the ability to interpret their meaning within an organisational context.
7.4 Provenance and Data Lineage
Provenance is another important component of discoverability. Users need to understand where information originated and how it has been transformed before it can be appropriately evaluated.
This connects directly with the discussion of accuracy in Chapter 5. Data lineage allows organisations to trace information through its transformation processes and determine which source systems contributed to a particular dataset. Qlik (2026) identifies lineage as an important mechanism for supporting data quality and governance.
The reproducibility literature provides further support for this perspective. Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes in machine learning. Similarly, Nature Computational Science (2021) highlights the importance of reproducibility in machine learning. Such documentation allows users to understand not only what data are available but also how those data came to exist in their current form.
For AI readiness, provenance therefore serves two related purposes. First, it supports evaluation, because users can assess whether the source and transformation history are appropriate. Second, it supports accountability, because organisations can establish how information entered an AI pipeline.
7.5 Discoverability and Reproducibility
Discoverability contributes directly to reproducibility because researchers and practitioners cannot reproduce an AI system if they cannot identify or reconstruct the data used in its development.
Heil et al. (2021) emphasise the importance of documenting data-related processes as part of reproducible machine learning. This suggests that discoverability should include sufficient information to establish which dataset or dataset version was used, where it originated and what transformations were applied.
The importance of this principle extends beyond academic research. In organisational AI systems, the ability to reconstruct the data foundations of a model may be necessary when investigating unexpected outputs, assessing changes in model performance or reviewing the appropriateness of previous decisions.
Discoverability therefore contributes to the auditability of AI systems by making relevant data and their associated documentation easier to identify and understand.
7.6 Human and Machine Discoverability
Discoverability has both human and machine dimensions.
Human users need interfaces and documentation that allow them to search for datasets, interpret metadata and evaluate suitability. Data scientists, engineers, business users and governance specialists may have different information requirements, but all require sufficient context to make informed decisions about data use.
Machine discoverability, by contrast, concerns whether computational systems can identify and interpret relevant information through structured metadata, semantic representations and machine-readable descriptions. This becomes increasingly important as AI systems themselves participate in information retrieval and data selection.
The distinction is significant because a dataset may be easy for a human to find but difficult for an automated system to interpret. Conversely, highly structured metadata may support automated retrieval while providing insufficient contextual explanation for human users. An effective AI-ready environment should therefore support both forms of discoverability.
7.7 Discoverability and Organisational Knowledge
Data discoverability is closely connected to the preservation of organisational knowledge. Data often contain information about business processes, customers, operations and historical decisions, but their meaning may depend on knowledge held by individuals or particular teams.
When contextual knowledge is not documented, the usefulness of the underlying data may decline as organisational structures and personnel change. Business glossaries, metadata and dataset documentation can therefore help convert implicit organisational knowledge into reusable information.
This is particularly important for AI because AI systems may require access to organisational knowledge that is distributed across databases, documents and other information sources. Without adequate documentation, an organisation may possess relevant information without possessing the contextual knowledge required to use it responsibly.
Discoverability can therefore be viewed as a mechanism through which organisations make their data assets and associated knowledge more reusable.
7.8 Discoverability and Data Governance
Discoverability also has a direct relationship with data governance. Janssen et al. (2020) argue that effective data governance is a foundation for trustworthy AI because organisations need structures through which data can be responsibly managed and used.
Governance becomes difficult when organisations cannot establish what data they possess, where those data originated, who is responsible for them or how they are being used. A data catalogue can therefore support governance by providing a structured inventory of organisational information assets.
Discoverability should consequently include governance-related metadata such as:
data ownership;
stewardship responsibilities;
classification and sensitivity;
permitted uses;
retention requirements;
provenance;
quality indicators; and
relationships to relevant AI applications.
These attributes enable organisations to connect data discovery with accountability. The objective is not merely to find data, but to establish the conditions under which those data can be appropriately used.
7.9 Discoverability and Security
As discussed in Chapter 6, discoverability must also be balanced with security. Making data easier to identify does not necessarily mean making their contents freely accessible.
This distinction supports the concept of controlled discoverability. Organisations may allow authorised users to identify that a dataset exists and obtain information about its purpose and characteristics while maintaining restrictions on access to the underlying data.
For example, metadata may communicate that a dataset contains sensitive information, identify its owner and explain its intended purpose without exposing the actual records. This allows discoverability to support governance without undermining security.
Qlik's (2026) positioning of discoverability alongside security is therefore significant. The two principles should be designed together so that increased visibility of the organisational data landscape is accompanied by appropriate controls over data access.
7.10 Discoverability and Data Quality
Discoverability also contributes to data quality because users are more likely to identify potential quality problems when appropriate documentation and contextual information are available.
A dataset with unclear definitions may be incorrectly interpreted, while information about collection methods and known limitations can help users assess whether apparent patterns are meaningful. Heil et al. (2021) emphasise the importance of documenting data collection and curation processes for reproducibility, supporting the broader argument that documentation is an essential component of responsible data use.
This creates a reinforcing relationship between discoverability and quality:
better documentation → better understanding → better data selection → more appropriate AI use.
Discoverability does not itself make data accurate, but it improves the ability of organisations to evaluate accuracy and suitability.
7.11 Discoverability in Generative AI and Large Language Models
Discoverability becomes particularly important in GenAI and LLM applications because these systems may need to locate relevant organisational information across large collections of documents and other data sources.
Qlik (2026) identifies retrieval-augmented generation (RAG) as an approach for connecting LLM applications with organisational information. In such environments, the effectiveness of retrieval depends partly on whether relevant information can be appropriately identified, indexed and interpreted.
Document metadata can therefore influence retrieval quality. Information concerning document type, subject, date, ownership and other contextual characteristics may help determine whether particular information is relevant to a user's query. Poorly documented or inconsistently structured information may be more difficult to retrieve appropriately.
This creates an important relationship between discoverability and consumability. Data must first be sufficiently discoverable and understandable before they can be transformed into forms that an AI architecture can effectively consume.
7.12 Documentation and AI Accountability
Documentation is also important for accountability. AI systems can produce outputs whose origins are difficult to explain if the underlying data environment is poorly documented. Although documentation cannot by itself provide complete explainability, it can establish an important evidential foundation for understanding what information was available to the system and how it was processed.
Heil et al. (2021) emphasise documentation as part of reproducible machine learning, while Janssen et al. (2020) highlight transparency and accountability as components of trustworthy AI governance. Together, these perspectives support the proposition that discoverability and documentation contribute to the organisational ability to scrutinise AI systems.
A well-documented data environment therefore provides more than convenience. It creates an evidential record through which organisations can evaluate the appropriateness of data use and investigate problems when they arise.
7.13 The Limits of Discoverability as an Independent Principle
Discoverability is essential, but it does not guarantee that data are suitable for AI. Data may be perfectly catalogued and extensively documented while remaining inaccurate, outdated, insecure or unrepresentative.
Similarly, increasing discoverability can create risks if sensitive information becomes easier to identify without appropriate access controls. Discoverability must therefore be evaluated alongside security and governance.
The principle should also not be confused with unrestricted accessibility. The objective is to make relevant information appropriately identifiable and understandable, not necessarily to make all information universally accessible.
This distinction is particularly important for organisational AI because responsible data use requires a balance between reuse and control.
7.14 Assessing Discoverability within AI Data Readiness
An AI-ready data environment should therefore provide mechanisms through which users and systems can:
identify relevant datasets and information sources;
understand the meaning and context of data;
access appropriate technical and business metadata;
establish ownership and stewardship;
trace provenance and lineage;
identify dataset versions and update histories;
understand known quality limitations;
determine appropriate and inappropriate uses;
distinguish sensitive from non-sensitive information; and
connect relevant data with AI applications and workflows.
These capabilities should be maintained throughout the data lifecycle. Metadata that are accurate when first created may become incomplete as datasets evolve, ownership changes or new uses emerge. Discoverability is therefore not a static catalogue-building exercise but an ongoing organisational capability.
7.15 Discoverability as an Interdependent Principle
The position of discoverability within Qlik's (2026) six-principle framework reinforces the interdependent nature of AI readiness. Discoverability supports accuracy because users can better evaluate data quality; it supports security because data can be classified and governed; it supports consumability because relevant information can be identified and prepared for AI architectures; and it supports timeliness because users can determine when information was collected or last updated.
At the same time, discoverability depends on the other principles. Data that are highly discoverable but insecure may create unacceptable risks. Data that are discoverable but inaccurate may be used incorrectly. Data that are discoverable but outdated may produce misleading AI outputs.
Discoverability should therefore be understood as an enabling capability that connects data assets with informed and governed use.
The central argument of this chapter is that AI-ready data must be more than technically accessible. They must be findable, interpretable, contextualised, traceable and appropriately governed. Metadata, semantic definitions, provenance and documentation provide the mechanisms through which this can be achieved.
Ultimately, discoverability transforms an organisation's data environment from a collection of isolated information assets into a more coherent and usable knowledge resource. For AI, this is particularly important because the value of data depends not only on their existence but on the ability of people and systems to identify relevant information, understand its meaning and determine whether it is appropriate for a particular application. Discoverability is therefore a critical bridge between data availability and responsible AI use, and an essential component of the broader AI data readiness framework.
8. Data Consumability for Machine Learning and Generative AI
Data consumability represents the sixth principle of AI-ready data and concerns the extent to which information can be effectively accessed, transformed and processed by an artificial intelligence (AI) system. Qlik (2026) identifies consumability as a key characteristic of AI-ready data and distinguishes between the requirements of traditional machine learning (ML) and generative artificial intelligence (GenAI) applications. This principle is important because the existence of data within an organisation does not necessarily mean that those data can be effectively used by an AI system. Data may be technically available but remain unusable because of incompatible formats, inadequate structure, insufficient metadata, poor semantic representation or inappropriate transformation processes.
Consumability therefore introduces an important distinction between data availability and data usability. Availability concerns whether information exists and can be accessed, whereas consumability concerns whether the information can be processed appropriately by the intended AI architecture. This distinction extends the concept of AI data readiness beyond conventional data quality. A dataset may be accurate, secure and discoverable but still require substantial transformation before it can support a particular ML model or GenAI application.
The importance of this distinction is consistent with the broader literature on dataset quality. Priestley et al. (2023) demonstrate that dataset quality in machine learning is multidimensional and extends across different stages of the data lifecycle. Data preparation and transformation are therefore not merely technical preprocessing activities; they form part of the conditions that determine whether information can be effectively incorporated into an AI system.
8.1 From Data Availability to Data Consumability
Data availability is a necessary starting point for AI, but it does not establish technical usability. Organisations frequently hold information in databases, documents, spreadsheets, applications and other systems that were designed for purposes other than machine learning or generative AI.
Such information may require transformation before it can be consumed effectively. Structured data may need to be cleaned and transformed into features appropriate for ML, while unstructured documents may require extraction, segmentation and semantic processing before they can be incorporated into GenAI applications.
Consumability should therefore be defined in relation to a specific AI purpose and architecture. There is no universal format that makes data consumable for every AI system. Data that are well suited to one model or application may require substantial modification for another.
This reinforces the contextual nature of AI readiness developed throughout this paper. Data should not be considered consumable in isolation; they are consumable for a particular technical purpose.
8.2 Consumability for Conventional Machine Learning
For conventional ML systems, consumability often involves transforming raw information into structured representations that can be processed by a model. This may include data cleaning, integration, feature construction and the preparation of training, validation and evaluation datasets.
The quality of these transformations is critical because preprocessing can affect the statistical characteristics of the data. A transformation that improves technical compatibility may also remove information, alter relationships between variables or introduce unintended assumptions.
Priestley et al. (2023) emphasise the importance of dataset quality throughout the machine learning lifecycle. This supports the view that consumability cannot be separated from data quality. Transformations should therefore be assessed not only according to whether they make data technically usable but also according to whether they preserve the characteristics required for reliable modelling.
Consumability for ML should consequently consider:
compatibility with the target model architecture;
appropriate data structures and formats;
preparation of training, validation and evaluation data;
suitable feature representations;
consistent preprocessing;
preservation of relevant information; and
documentation of transformation processes.
These requirements help ensure that technical preparation does not undermine the original meaning or quality of the information.
8.3 Consumability for Generative AI
GenAI introduces different requirements because applications may process large volumes of unstructured or semi-structured information. Rather than relying exclusively on structured numerical features, GenAI systems may need access to documents, organisational knowledge and other forms of contextual information.
Qlik (2026) identifies retrieval-augmented generation (RAG) as an approach for connecting large language models (LLMs) with organisational information. In such architectures, data must be prepared in ways that support effective retrieval and subsequent processing by the LLM.
This may require extracting information from documents, dividing content into meaningful units, generating representations suitable for semantic retrieval and associating relevant metadata with the resulting information. The objective is not simply to convert documents into a machine-readable format but to preserve sufficient context and meaning for the intended application.
Consumability in GenAI should therefore be understood as involving both technical compatibility and semantic suitability.
8.4 Retrieval-Augmented Generation and the Data Foundation
RAG provides a useful example of why consumability is broader than format compatibility. A RAG system typically depends on an external knowledge source that can be searched or retrieved in response to a user's query. The quality of the resulting response therefore depends partly on the quality of the information made available to the retrieval process.
Qlik (2026) identifies RAG as a mechanism for incorporating organisational information into LLM applications. This creates a direct relationship between consumability and the other AI-readiness principles.
For retrieval to be effective, information must be:
sufficiently accurate;
appropriately current;
relevant to the intended use;
discoverable within the knowledge environment;
represented in a form suitable for retrieval; and
appropriately governed and secured.
Consequently, consumability cannot compensate for weaknesses elsewhere in the data environment. A technically sophisticated retrieval architecture may still produce poor outputs if the underlying information is inaccurate, outdated or poorly documented.
8.5 Data Transformation and Preservation of Meaning
Transformation is central to consumability but introduces a potential risk: making data technically usable can alter their meaning.
Raw information may contain contextual relationships that are not preserved when data are reformatted, aggregated, segmented or otherwise transformed. For example, converting complex source information into simplified representations may improve computational efficiency while removing information that is relevant to interpretation.
This creates an important requirement for AI-ready data: transformations should be evaluated for their impact on both technical usability and semantic fidelity.
Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes for reproducible machine learning. This principle is directly relevant to consumability because users need to understand how source information was transformed before being incorporated into an AI system.
Documentation should therefore establish:
the original source of the information;
the transformations applied;
the reasons for those transformations;
the resulting representation;
known limitations introduced by processing; and
the relationship between transformed data and the original source.
Such documentation provides a bridge between consumability, provenance and reproducibility.
8.6 Structured and Unstructured Data
The distinction between structured and unstructured information is particularly relevant to AI consumability. Traditional ML applications often operate on structured datasets in which variables and observations have predefined formats. GenAI applications, by contrast, frequently need to process unstructured information such as documents and text.
This does not mean that structured information is inherently more consumable. Structured data can contain poorly defined variables, inconsistent semantics or inappropriate representations, while unstructured information may contain rich contextual knowledge that can become highly useful after appropriate processing.
The relevant question is therefore whether the information has been transformed into a representation appropriate for the intended AI task while retaining the meaning required for reliable use.
8.7 Metadata and Semantic Representation
Consumability is closely connected to discoverability because AI systems require more than raw data to determine how information should be interpreted. Metadata can provide information about the source, meaning, date, ownership and other characteristics of data.
Qlik (2026) identifies metadata and semantic capabilities as important components of AI-ready data. The relevance of these capabilities extends into consumability because metadata can provide contextual information required by AI pipelines and retrieval systems.
Semantic representation is particularly important for GenAI applications. Information that is technically retrievable but lacks sufficient context may be interpreted incorrectly or provide incomplete support for generated responses. Consumability therefore requires a balance between computational representation and preservation of semantic meaning.
This creates a direct relationship between Chapters 7 and 8: discoverability helps establish what information exists and what it means, while consumability concerns whether that information can be effectively transformed and used by the target AI system.
8.8 Interoperability and Standardisation
Interoperability is another important component of consumability. Organisations frequently operate multiple systems using different formats, structures and conventions. Data may therefore require integration or transformation before they can be incorporated into a common AI pipeline.
Standardised and interoperable representations can reduce unnecessary transformation and support reuse across different AI applications. However, standardisation should not be pursued at the expense of meaning or contextual information.
Priestley et al. (2023) emphasise the multidimensional nature of dataset quality, supporting the need to consider the consequences of transformation and integration rather than treating technical standardisation as an end in itself.
Consumability should therefore be evaluated according to whether data can be moved and transformed between relevant systems without unacceptable loss of quality, meaning or provenance.
8.9 Consumability and Data Quality
Consumability is fundamentally dependent on data quality. Poor-quality information may become easier to process after transformation, but technical usability does not remove underlying inaccuracies or biases.
For example, converting an inaccurate document collection into a highly efficient retrieval system does not make the underlying information reliable. Similarly, transforming incomplete data into model-ready features does not resolve the consequences of missing information.
Priestley et al. (2023) provide a basis for viewing these issues through the broader concept of dataset quality. Consumability should therefore be treated as complementary to accuracy rather than as an alternative to it.
This distinction is particularly important because highly automated AI pipelines may make poor-quality information easier to process at scale. The greater the efficiency of the pipeline, the greater the potential for underlying data problems to propagate rapidly through AI applications.
8.10 Consumability, Data Lineage and Reproducibility
The transformation processes associated with consumability also create a need for lineage and provenance. Organisations should be able to determine how the data consumed by an AI system relate to their original sources.
Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes to support reproducible machine learning. The same principle applies to AI consumability: transformations should be recorded sufficiently to enable users to understand and, where necessary, reconstruct how source information was converted into AI-ready representations.
This is particularly important when AI outputs need to be investigated. If a generated response or model prediction appears unreliable, organisations may need to establish whether the problem originated in the source information, the transformation process, the retrieval mechanism or the AI model itself.
Consumability therefore creates an important link between technical preparation and accountability.
8.11 Consumability and Security
The transformation of data can also create new security considerations. Data may pass through multiple processing environments before becoming consumable by an AI system. Each additional processing stage may introduce new locations where sensitive information is stored or accessed.
This connects directly with the security principle discussed in Chapter 6. Data cannot be considered fully AI-ready merely because they are technically consumable if the transformation process creates inappropriate exposure of sensitive information.
In GenAI environments, this issue becomes particularly important when organisational documents are indexed or transformed for retrieval. Security controls must remain effective throughout the transformation and retrieval processes rather than applying only to the original source repository.
Consumability should therefore be achieved in a way that preserves appropriate access restrictions and governance requirements.
8.12 Consumability and Timeliness
Consumability also interacts with timeliness. Data may be technically prepared for an AI application but become less useful when the underlying information changes.
This is particularly relevant to RAG systems, where organisational knowledge may be updated independently of the underlying LLM. Qlik (2026) highlights RAG as a means of incorporating organisational information into LLM applications, making the maintenance of the associated knowledge environment an important component of operational AI readiness.
The technical ability to process updated information must therefore be considered alongside the frequency with which that information changes. A system that can consume data efficiently but cannot accommodate updates may become increasingly disconnected from its operating environment.
8.13 Consumability and Diversity
The relationship between consumability and diversity is also important. Transforming data into a standardised computational representation may inadvertently reduce the visibility of less common groups, contexts or forms of information.
This reinforces the principle established in Chapter 3 that data preparation should not be evaluated solely according to technical efficiency. Organisations should consider whether transformations preserve the relevant diversity and representational characteristics of the source data.
The objective should therefore be to make diverse information technically usable without erasing meaningful variation.
8.14 The Limits of Consumability as an Independent Principle
Consumability is essential, but it cannot establish AI readiness by itself. Highly consumable data may still be inaccurate, outdated, insecure, unrepresentative or poorly documented.
There is also a risk that organisations may prioritise technical convenience over data suitability. Data that are easy to ingest may be preferred because they require less engineering effort, even when alternative information would provide a more appropriate representation of the problem.
This creates a potential bias towards what is technically convenient rather than what is substantively appropriate. AI readiness should therefore require organisations to evaluate whether the data are appropriate for the AI task before investing in the technical processes required to consume them.
Consumability should consequently be viewed as an enabling property rather than the ultimate measure of data quality.
8.15 Assessing Consumability within AI Data Readiness
An AI-ready data environment should assess whether data can be:
accessed by authorised AI systems;
represented in formats compatible with the target architecture;
transformed without unacceptable loss of meaning;
integrated across relevant systems;
appropriately structured for ML applications;
prepared for semantic retrieval in GenAI applications;
accompanied by sufficient metadata and contextual information;
traced back to their original sources;
updated as the underlying information changes; and
processed without undermining security, governance or data quality.
These requirements demonstrate that consumability is both technical and contextual. The relevant question is not simply whether data can be processed, but whether they can be processed appropriately for the intended AI application.
8.16 Consumability as an Interdependent Principle
The position of consumability as the sixth principle in Qlik's (2026) framework provides a useful culmination of the preceding dimensions. Data must be diverse enough to represent relevant conditions, timely enough to remain useful, accurate enough to support reliable outputs, secure enough to permit responsible use and discoverable enough to be identified and understood. Consumability provides the mechanism through which these characteristics can be translated into practical use within AI architectures.
At the same time, consumability depends on the other principles. Accurate information that cannot be technically processed remains unavailable in practice. Highly consumable information that is inaccurate or insecure may simply allow poor or risky information to be processed more efficiently. Consumability therefore cannot be considered independently from the wider data-readiness framework.
The interaction can be summarised conceptually as follows:
Discoverability identifies and contextualises the information → transformation makes it technically usable → consumability enables its integration into the AI system → monitoring ensures that its suitability is maintained.
This sequence should not, however, be interpreted as strictly linear. Changes in AI architecture may create new consumability requirements, while changes in data quality or governance may require transformations to be reconsidered.
8.17 Consumability as an Organisational Capability
The development of consumable data also requires organisational capabilities beyond technical engineering. Data engineers, data scientists, domain experts, governance specialists and security professionals may need to collaborate to determine how information should be transformed and what characteristics must be preserved.
This multidisciplinary requirement reflects the broader governance perspective of Janssen et al. (2020), who position data governance as a foundation for trustworthy AI. Consumability should therefore be incorporated into organisational data governance rather than treated solely as an engineering problem.
Clear ownership is particularly important where transformations affect the meaning or suitability of information. Organisations should establish who determines the appropriate representation, who validates the resulting data and who is responsible for maintaining the pipeline as requirements change.
8.18 Towards a Definition of AI Data Consumability
Based on the preceding analysis, AI data consumability can be defined as:
The capacity of data to be appropriately accessed, interpreted, transformed and processed by a specified AI architecture while preserving sufficient quality, meaning, provenance, security and contextual relevance for the intended application.
This definition is broader than technical compatibility. It recognises that data may be syntactically processable while remaining semantically inappropriate or organisationally unsafe.
It also reinforces the contextual nature of AI readiness. Consumability cannot be established independently of an intended use because the requirements of a conventional ML model may differ substantially from those of an LLM-based RAG application.
8.19 Conclusion
The analysis of consumability demonstrates that AI readiness requires more than data availability. Data must be capable of being transformed and incorporated into AI architectures in ways that preserve their quality, meaning, provenance and appropriate governance.
For conventional ML, this involves preparing structured representations and features that are compatible with the target modelling task. For GenAI and LLM applications, it may involve transforming unstructured information into representations that support semantic retrieval and contextual generation. In both cases, transformation introduces risks that must be managed through documentation, lineage and validation (Heil et al., 2021).
Qlik's (2026) consumability principle therefore provides an important final dimension of the AI-ready data framework, but its significance lies in its relationship with the other principles. Data cannot be effectively consumed if they are inaccurate, outdated, insecure or undiscoverable. Conversely, the other dimensions of readiness have limited operational value if information cannot ultimately be incorporated into the intended AI architecture.
The central argument is therefore that AI data consumability is the capacity to make information technically usable without compromising its substantive quality, meaning, provenance or governance. This interpretation moves consumability beyond simple format compatibility and establishes it as an essential bridge between organisational data assets and operational AI systems.
Ultimately, consumability represents the point at which AI data readiness becomes operational. It connects the quality and governance characteristics established in the preceding principles with the technical architectures through which AI systems actually use information. For this reason, consumability should be assessed continuously and in relation to the specific AI application, rather than treated as a one-time technical preparation exercise.
9. AI Readiness as a Multidimensional and Dynamic Framework
The preceding chapters have examined six principles of AI-ready data: diversity, timeliness, accuracy, security, discoverability and consumability. While each dimension represents an important characteristic of data suitable for artificial intelligence (AI), treating them as independent criteria would provide an incomplete understanding of AI readiness. The central proposition of this paper is that AI readiness is a multidimensional and dynamic property arising from the interaction between these dimensions rather than from the achievement of any single characteristic.
Qlik (2026) provides a practical framework in which the six principles collectively establish a trusted data foundation for AI. The academic literature considered in this paper supports the underlying logic of this approach. Priestley et al. (2023) demonstrate that dataset quality in machine learning (ML) consists of multiple dimensions and must be considered across the data lifecycle. Janssen et al. (2020) similarly position data governance as a foundation for trustworthy AI, while Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes for reproducible machine learning. These perspectives suggest that AI readiness should not be reduced to a single technical attribute or one-off assessment.
Instead, AI readiness should be understood as the extent to which data remain appropriate, reliable, governed and technically usable for a specified AI purpose over time.
9.1 From Six Principles to an Integrated Framework
The six principles provide complementary perspectives on the suitability of data for AI.
Diversity concerns whether data adequately capture the relevant populations, conditions, perspectives and patterns required by the intended application.
Timeliness concerns whether information remains sufficiently current and temporally relevant to the environment in which the AI system operates.
Accuracy concerns whether data are sufficiently correct, complete, consistent and reliable for their intended purpose.
Security concerns whether information is appropriately protected and governed throughout its lifecycle.
Discoverability concerns whether relevant information can be identified, interpreted, contextualised and traced.
Consumability concerns whether data can be appropriately accessed, transformed and processed by the target AI architecture without unacceptable loss of meaning, quality or provenance.
These dimensions should not be interpreted as six independent stages. Instead, they form an interconnected system in which weaknesses in one dimension can affect the effectiveness of others.
For example, highly diverse data may improve representativeness but provide limited value if they are inaccurate. Accurate data may become unsuitable when they are significantly outdated. Discoverable information may create unacceptable risks if security controls are inadequate. Similarly, data that are technically consumable may still produce unreliable AI outputs if their provenance or contextual meaning is unclear.
AI readiness is therefore best conceptualised as a system property.
9.2 The Interdependence of the Six Dimensions
The interaction between the dimensions can be illustrated through several relationships.
First, diversity and accuracy are complementary. A dataset may contain highly accurate observations but still provide an incomplete representation of the population or environment. Bernhardt, Jones and Glocker (2022) demonstrate that dataset bias can arise through multiple mechanisms, illustrating why accurate individual observations do not necessarily guarantee representative AI data.
Second, timeliness and accuracy are closely connected. Information can be factually accurate at the point of collection while becoming less relevant as the environment changes. AI readiness must therefore consider not only whether data are correct but whether they remain appropriate for the decision context.
Third, security and discoverability must be balanced. Organisations need to know what data they possess and understand their characteristics, but increased visibility must not result in inappropriate access to sensitive information. This supports the concept of controlled discoverability developed in Chapter 7.
Fourth, discoverability and consumability are mutually reinforcing. Information must be identifiable and understandable before it can be appropriately transformed for an AI system. Metadata, semantic definitions and provenance can therefore support both discovery and technical preparation.
Finally, consumability and accuracy must remain connected. Transformation may make information technically usable while introducing errors or removing important contextual information. Consumability therefore requires validation that transformed representations remain sufficiently faithful to their sources.
These relationships demonstrate why AI readiness cannot be represented adequately by a collection of unrelated data-quality indicators.
9.3 AI Readiness as a Context-Dependent Property
An important implication of the integrated framework is that AI readiness is context-dependent. There is no universal threshold at which a dataset becomes AI-ready for every possible application.
The requirements for an AI system used for exploratory analysis may differ from those for a system supporting consequential organisational decisions. Similarly, the requirements for a conventional predictive model may differ from those for a generative AI application using retrieval-augmented generation (RAG).
Priestley et al. (2023) highlight the importance of evaluating dataset quality in relation to machine learning requirements. This supports the broader argument that data suitability must be assessed against the purpose for which the data are being used.
Consequently, AI readiness should be expressed in relation to at least three contextual factors:
The data – what information is available, how it was collected and what characteristics it possesses.
The AI application – what the system is intended to do and what data requirements follow from that purpose.
The operating environment – the population, organisational context and conditions in which the system will be deployed.
This means that the same dataset could be considered AI-ready for one application but unsuitable for another.
9.4 AI Readiness as a Lifecycle Property
AI readiness should also be understood as dynamic rather than static. A dataset that is suitable at one point in time may become unsuitable as the data environment, organisational requirements or AI architecture changes.
This is particularly evident in relation to timeliness. Changes in customer behaviour, organisational processes, external conditions or other environmental factors can alter the relevance of previously collected information. Consequently, AI readiness must be continuously reassessed.
The lifecycle perspective is consistent with Priestley et al. (2023), who conceptualise dataset quality as an issue extending across multiple stages of machine learning. Heil et al. (2021) similarly emphasise the importance of documenting data processes rather than treating datasets as isolated objects.
The implication is that AI readiness should be incorporated into ongoing data management rather than performed only during initial AI project development.
A simplified lifecycle can be conceptualised as:
Data acquisition → assessment → preparation → AI use → monitoring → reassessment → updating or remediation
At each stage, the six dimensions may change.
9.5 The Role of Monitoring and Continuous Assessment
Continuous assessment is particularly important because data quality problems may emerge after deployment. Changes in source systems, data collection procedures, population characteristics or organisational processes can affect the suitability of information.
Monitoring should therefore consider both the condition of the data and the environment in which they are being used.
For example:
diversity monitoring can identify changes in representation;
timeliness monitoring can identify outdated information;
accuracy monitoring can detect emerging quality problems;
security monitoring can identify inappropriate access;
discoverability monitoring can identify incomplete or obsolete metadata; and
consumability monitoring can identify failures in AI data pipelines.
This suggests that AI readiness should be treated as a continuously monitored state, rather than a certification that remains valid indefinitely.
9.6 Thresholds and Maturity
One of the challenges associated with the six-dimensional framework is determining how readiness should be measured. Qlik (2026) proposes an AI Trust Score that aggregates indicators associated with the six principles into a composite measure.
A composite score has practical advantages because it can provide organisations with a simple way to summarise progress and compare data environments. However, a single aggregate measure can also conceal important weaknesses.
For example, a dataset might achieve high scores for diversity, discoverability and consumability while failing to meet minimum security requirements. A high overall score should not compensate for a critical weakness that makes the data inappropriate for use.
This suggests a distinction between threshold requirements and maturity measures.
Threshold requirements establish minimum conditions that must be satisfied before data can be used for a particular AI application. Maturity measures, by contrast, indicate the extent to which an organisation has developed more advanced capabilities.
An organisation could therefore assess AI readiness through two complementary questions:
Is the data sufficiently safe and appropriate to use?
How mature is the organisation's capability to manage and improve AI-ready data?
This approach provides greater analytical value than relying exclusively on a single numerical score.
9.7 A Three-Level AI Data Readiness Model
Building on the six principles, AI readiness can be conceptualised through three broad levels.
Level 1: Foundational Readiness
At the foundational level, data meet minimum requirements for responsible use. This includes appropriate security, governance, basic quality controls and sufficient documentation.
The primary question at this level is:
Can the organisation responsibly use these data for the intended AI purpose?
Failure to meet critical requirements at this level should prevent deployment regardless of performance in other dimensions.
Level 2: Operational Readiness
At the operational level, data demonstrate sufficient accuracy, timeliness, diversity, discoverability and technical consumability for the intended AI application.
The focus shifts from basic permission to effective operational use.
The primary question becomes:
Can these data reliably support the intended AI system?
Level 3: Adaptive Readiness
At the adaptive level, organisations continuously monitor and improve their data environments in response to changes in AI requirements, data sources and operating conditions.
The primary question becomes:
Can the organisation maintain data suitability as circumstances change?
This third level is particularly important because AI systems increasingly operate in environments characterised by continuous data generation and changing organisational requirements.
9.8 AI Readiness and Trustworthy AI
The concept of AI readiness is closely connected to trustworthy AI. Janssen et al. (2020) argue that data governance is fundamental to organising data for trustworthy AI. Their perspective supports the proposition that technical performance alone is insufficient to establish trustworthiness.
An AI system may produce accurate predictions while still raising concerns about privacy, representation, transparency or accountability. The six-principle framework addresses these concerns by extending attention beyond accuracy towards security, diversity, discoverability and governance-related characteristics.
This does not mean that AI-ready data automatically guarantee trustworthy AI. Model architecture, implementation, human oversight and organisational decision-making also influence whether an AI system is trustworthy. However, inadequate data foundations can undermine these broader efforts.
AI data readiness should therefore be understood as a necessary but not sufficient condition for trustworthy AI.
9.9 AI Readiness and Reproducibility
Reproducibility provides another important perspective on the framework. Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation alongside machine learning models and algorithms. Nature Computational Science (2021) similarly identifies reproducibility as an important concern for machine learning.
This supports the inclusion of discoverability, provenance and consumability within the AI readiness framework. If organisations cannot establish what data were used, where those data came from or how they were transformed, it becomes difficult to reproduce or investigate an AI system.
Reproducibility therefore provides an additional rationale for treating AI readiness as a lifecycle capability rather than a property of a static dataset.
9.10 AI Readiness and Algorithmic Bias
The framework also provides a more comprehensive way of considering algorithmic bias. Bernhardt, Jones and Glocker (2022) demonstrate that dataset bias can arise through multiple and interacting mechanisms. This illustrates why diversity cannot be treated as a simple numerical measure of the number of groups or sources represented in a dataset.
AI readiness requires organisations to consider whether data adequately represent the intended deployment environment and whether collection and measurement processes introduce systematic limitations.
The interaction between diversity and the other dimensions is also important. For example, information about underrepresented groups may exist but remain undiscoverable, outdated or technically difficult to incorporate into an AI pipeline. Conversely, a highly consumable dataset may encode systematic representational limitations.
The six-dimensional framework therefore encourages a more holistic approach to bias by connecting representation with data quality, context, governance and technical use.
9.11 AI Readiness and Organisational Capability
The framework has implications beyond individual datasets. Organisations require capabilities that enable them to create, maintain and govern AI-ready data environments.
Janssen et al. (2020) emphasise organisational arrangements for data governance, while Heil et al. (2021) highlight the importance of systematic documentation. Together, these perspectives indicate that AI readiness depends partly on organisational processes and responsibilities.
Relevant organisational capabilities include:
data ownership and stewardship;
quality management;
metadata and catalogue management;
lineage and provenance tracking;
security and access management;
data engineering and transformation;
monitoring and assessment; and
cross-functional governance.
AI readiness should therefore be viewed partly as an organisational maturity issue. Organisations that repeatedly develop AI applications without establishing these supporting capabilities may encounter recurring data problems, even when individual projects appear technically successful.
9.12 A Proposed AI Data Readiness Matrix
The six principles can be operationalised through a readiness assessment in which each dimension is evaluated in relation to the intended AI application and its operating environment. Diversity requires organisations to determine whether the data adequately represent the relevant population, circumstances and operational conditions, with particular attention to variation, underrepresented groups and potential sampling bias. Timeliness concerns whether the data are sufficiently current and temporally relevant for the intended application, taking into account data freshness, update frequency, temporal coverage and changes in the operating environment. Accuracy requires assessment of whether the data are sufficiently reliable for their intended purpose, including their correctness, completeness and consistency, as well as the presence of anomalies and the availability of appropriate data lineage.
Security concerns whether the data are appropriately protected and governed throughout their lifecycle. This includes consideration of data classification, access controls, integrity, privacy and monitoring of data use. Discoverability addresses whether users and AI systems can identify, locate and understand relevant data. Effective discoverability depends on appropriate metadata, semantic definitions, provenance, catalogues and documentation. Finally, consumability concerns whether the data can be appropriately accessed, transformed and used by the target AI architecture. This includes considerations such as data formats, transformation processes, semantic representation, interoperability and validation.
These dimensions should not be interpreted as a checklist in which every criterion has equal importance in every context. Rather, the assessment should reflect the requirements, risks and objectives of the specific AI application. The framework provides a structured basis for identifying strengths and weaknesses in an organisation's data environment, while also revealing dependencies between the dimensions. For example, data that are highly consumable but inaccurate may undermine model reliability, while highly discoverable data that lack adequate security controls may create privacy and governance risks. AI readiness should therefore be assessed holistically, with particular attention given to critical dependencies and application-specific requirements.
9.13 From Assessment to Decision-Making
The purpose of an AI readiness assessment should ultimately be to support better decisions about data and AI deployment.
An assessment could result in one of four broad outcomes:
Ready: The data meet the required thresholds for the intended AI application.
Conditionally ready: The data can be used provided that specified weaknesses are addressed or appropriate controls are implemented.
Not ready: Critical deficiencies prevent responsible or reliable use.
Under review: The data require further investigation because their suitability cannot yet be established.
This approach is preferable to treating AI readiness as a binary property because it allows organisations to identify specific remediation requirements.
For example, data may be technically consumable but require improvements in security before deployment. Alternatively, a highly secure dataset may require additional documentation and metadata before it can be responsibly incorporated into a GenAI application.
9.14 AI Readiness as a Dynamic Feedback System
The framework can be further understood as a feedback system rather than a linear process.
Initial assessment establishes whether data are suitable for an intended AI application. Once deployed, monitoring generates information about changes in data quality, relevance and usage. Those observations can then trigger remediation, transformation, additional governance or changes to the AI application.
The process can therefore be represented conceptually as:
Assess → Prepare → Deploy → Monitor → Learn → Remediate → Reassess
This feedback loop is particularly important because AI systems may operate in environments where requirements evolve. Data readiness must consequently evolve with them.
This perspective also reinforces the argument that AI readiness is not a one-time certification. It is an ongoing capability for maintaining the suitability of data in relation to changing AI requirements.
9.15 Limitations of a Composite AI Readiness Score
Although a composite score such as the AI Trust Score proposed by Qlik (2026) may be useful for communication and high-level management, it should not become a substitute for multidimensional assessment.
Aggregation creates the possibility of compensatory scoring, whereby strengths in some dimensions offset serious weaknesses in others. Such compensation may be inappropriate when particular dimensions represent non-negotiable requirements.
For example, strong discoverability cannot compensate for inadequate security where sensitive information is involved. Similarly, high consumability cannot compensate for severe inaccuracies in the underlying data.
A more defensible approach would therefore combine:
minimum thresholds for critical dimensions;
dimension-specific scores for detailed assessment;
overall maturity indicators for organisational development; and
qualitative assessment of contextual risks and limitations.
This approach preserves the practical value of measurement while avoiding the false precision that can arise from reducing a complex data environment to a single number.
9.16 Towards an Integrated Definition of AI Data Readiness
The analysis across Chapters 2–8 supports a broader definition of AI data readiness.
AI data readiness can be understood as:
The multidimensional and continuously maintained capacity of data to support a specified AI application by being sufficiently diverse, timely, accurate, secure, discoverable and consumable within its organisational and operational context.
This definition contains several important elements.
First, AI readiness is multidimensional because no single data characteristic is sufficient.
Second, it is contextual because suitability depends on the intended AI application and operating environment.
Third, it is dynamic because data and their surrounding environments change over time.
Fourth, it is organisational because maintaining readiness requires governance, documentation, monitoring and accountability.
Finally, it is application-oriented because data are not simply ready or unready in the abstract; they are ready or unready for a particular purpose.
9.17 Theoretical Contribution of the Framework
The principal contribution of the framework is to position AI data readiness between traditional data quality and broader trustworthy AI governance.
Traditional data quality remains essential, but AI introduces additional requirements concerning representation, context, technical usability and governance. Priestley et al. (2023) provide a foundation for understanding the multidimensional nature of dataset quality, while Janssen et al. (2020) demonstrate the importance of governance for trustworthy AI. Heil et al. (2021) and Nature Computational Science (2021) further establish the importance of documentation and reproducibility.
The six-principle framework integrates these concerns into a single conceptual structure. It therefore provides a way of connecting data quality, data governance and AI engineering rather than treating them as separate organisational domains.
The framework does not claim that the six dimensions encompass every aspect of trustworthy AI. Rather, it provides a data-centred foundation from which broader AI governance and system-level assessments can be developed.
9.18 Practical Significance for Organisations
For organisations, the framework suggests that AI initiatives should begin with a structured assessment of the data foundation rather than with model selection alone.
Before deploying an AI application, organisations should establish:
whether the data adequately represent the intended environment;
whether the information remains sufficiently current;
whether its quality is adequate for the intended purpose;
whether security and privacy controls are appropriate;
whether relevant information can be discovered and understood; and
whether the information can be reliably consumed by the target AI architecture.
Where deficiencies are identified, organisations should determine whether they can be remediated through additional data collection, quality improvement, governance, transformation or architectural changes.
This approach can reduce the risk of developing technically sophisticated AI systems on inadequate data foundations.
9.19 Conclusion
The six principles of diversity, timeliness, accuracy, security, discoverability and consumability provide a useful foundation for conceptualising AI data readiness. However, their greatest value emerges when they are considered as an integrated system rather than as independent characteristics.
The analysis demonstrates that AI readiness is multidimensional, contextual, lifecycle-based and organisational. Diversity addresses representation; timeliness addresses temporal relevance; accuracy addresses reliability; security addresses protection and responsible access; discoverability addresses understanding and reuse; and consumability addresses technical and semantic usability. None of these dimensions is sufficient on its own.
The framework also demonstrates the limitations of treating AI readiness as a single numerical score. Although composite measures such as Qlik's (2026) AI Trust Score may provide useful high-level indicators, critical weaknesses should not necessarily be offset by strengths elsewhere. A more robust approach combines minimum thresholds, dimension-specific assessment and continuous organisational maturity measurement.
The resulting framework positions AI readiness as a continuous capability rather than a one-time property of a dataset. Organisations must assess, prepare, deploy, monitor and reassess their data as AI applications and operating environments evolve. This lifecycle perspective is consistent with research emphasising dataset quality, governance and reproducibility (Priestley et al., 2023; Heil et al., 2021; Janssen et al., 2020).
Ultimately, AI readiness provides a bridge between data management and trustworthy AI. It recognises that effective AI depends not simply on the availability of large quantities of information but on whether that information is appropriate, representative, current, reliable, secure, understandable and technically usable. The six-principle framework therefore provides a foundation for organisations seeking to transform their data assets into sustainable and responsible AI capabilities.
10. Implications for Organisations
The preceding analysis establishes that AI data readiness is not simply a technical characteristic of a dataset but an organisational capability that must be developed, governed and maintained over time. The six principles of diversity, timeliness, accuracy, security, discoverability and consumability provide a structured basis for assessing whether organisational data are suitable for artificial intelligence (AI) applications (Qlik, 2026). However, translating these principles into practice requires changes in governance, processes, technical infrastructure, organisational responsibilities and ongoing monitoring.
The implications are particularly significant as organisations expand their use of machine learning (ML), generative artificial intelligence (GenAI) and large language model (LLM) applications. AI initiatives can create substantial value when supported by appropriate data, but weaknesses in the underlying data environment can reduce system reliability, increase operational risk and undermine trust. Priestley et al. (2023) emphasise the multidimensional nature of dataset quality in machine learning, while Janssen et al. (2020) identify data governance as a foundation for trustworthy AI. These perspectives suggest that organisations should treat data readiness as a strategic capability rather than as a preparatory task undertaken immediately before model development.
10.1 Moving from AI Projects to AI Data Capability
A common organisational approach is to treat AI as a sequence of individual projects: identify a business problem, select a model, obtain data and develop an application. Although this approach may be appropriate for experimentation, it becomes increasingly problematic as AI adoption expands across an organisation.
Repeated AI projects may encounter the same underlying data problems, including fragmented information, inconsistent definitions, poor documentation, outdated datasets and inadequate governance. Addressing these issues separately within each project can result in duplicated effort and inconsistent standards.
A more sustainable approach is to develop an organisational AI data capability. This capability should provide reusable mechanisms for identifying, assessing, preparing, governing and monitoring data across multiple AI applications.
Qlik (2026) provides a practical foundation for such a capability through the six principles of AI-ready data. The organisational challenge is therefore to translate these principles into policies, processes, technologies and responsibilities that can operate consistently across the AI lifecycle.
10.2 Integrating Data Governance with the AI Lifecycle
The first major implication is that data governance should be integrated into the AI lifecycle rather than treated as a separate administrative activity.
Janssen et al. (2020) argue that data governance is fundamental to organising data for trustworthy AI. Governance should therefore begin when an AI use case is being defined and continue through data acquisition, preparation, development, deployment and monitoring.
An integrated lifecycle could include:
Use-case definition → data identification → readiness assessment → data preparation → model or application development → validation → deployment → monitoring → reassessment
At each stage, organisations should evaluate whether the six dimensions of AI readiness remain appropriate.
For example, during use-case definition, organisations should establish what data are required and what level of diversity or timeliness is necessary. During preparation, they should assess accuracy, lineage and consumability. During deployment, security and access controls become particularly important. Following deployment, monitoring should identify whether changing conditions have affected the continuing suitability of the data.
This approach prevents data governance from becoming a final compliance checkpoint and instead makes it part of AI development itself.
10.3 Establishing Clear Data Ownership and Accountability
Effective AI data management requires clear responsibility for data assets. Organisations should establish who owns, manages and maintains important datasets and who is accountable for decisions concerning their use.
Janssen et al. (2020) emphasise organisational arrangements as an important component of data governance. This supports the need for clearly defined roles rather than assuming that responsibility for data quality belongs exclusively to technical teams.
Depending on organisational structure, responsibilities may include:
Data owners, responsible for the organisational purpose and appropriate use of data;
Data stewards, responsible for definitions, quality and metadata;
Data engineers, responsible for pipelines, transformation and technical integration;
Data scientists, responsible for analytical and modelling requirements;
Security specialists, responsible for appropriate protection and access controls; and
Domain experts, responsible for contextual interpretation and validation.
These roles should not operate independently. AI data readiness requires collaboration between technical, governance and domain perspectives.
10.4 Establishing AI Data Readiness Assessment
Organisations should introduce a structured assessment process before data are incorporated into significant AI applications.
The assessment should examine the six dimensions identified by Qlik (2026):
Diversity – Are the data sufficiently representative of the intended population, context and operating environment?
Timeliness – Are the data sufficiently current and temporally relevant?
Accuracy – Are the data sufficiently correct, complete and consistent for the intended purpose?
Security – Are appropriate privacy, access and protection controls in place?
Discoverability – Can users and systems identify, understand and trace the information?
Consumability – Can the information be appropriately transformed and processed by the target AI architecture?
The assessment should be contextual rather than purely numerical. Different AI applications will place different levels of importance on particular dimensions.
For example, a system operating in a rapidly changing environment may require stringent timeliness controls, while a system using sensitive personal information may require particularly strong security and governance requirements.
10.5 Establishing Minimum Readiness Thresholds
Chapter 9 established that AI readiness should not rely exclusively on a single composite score. Organisations should therefore establish minimum thresholds for dimensions where failure could create unacceptable risk.
Qlik (2026) proposes an AI Trust Score as a means of aggregating indicators associated with AI-ready data. Such an approach may be useful for high-level monitoring, but organisational decision-making should also recognise that certain weaknesses cannot reasonably be compensated for by strengths elsewhere.
For example, high levels of data diversity and discoverability should not compensate for inadequate security where sensitive information is involved. Similarly, highly consumable data should not be approved for an AI application if their accuracy is insufficient for the intended decision.
Organisations should therefore distinguish between:
mandatory requirements, which must be satisfied before deployment;
risk-based requirements, which depend on the application and data context; and
maturity indicators, which measure opportunities for continuous improvement.
This approach provides greater flexibility than a simple pass-or-fail assessment while maintaining appropriate safeguards.
10.6 Investing in Metadata and Data Documentation
One of the most important organisational investments is the development of metadata and documentation capabilities.
Discoverability requires more than searchable storage. Users need information about the meaning, origin, ownership, quality, limitations and appropriate use of datasets. Qlik (2026) identifies metadata, business glossaries and catalogues as important components of AI-ready data.
Documentation also supports reproducibility. Heil et al. (2021) emphasise the importance of documenting data collection, cleaning and curation processes in machine learning. Such documentation enables organisations to understand how data entered an AI workflow and how they were transformed.
Organisations should therefore maintain documentation covering, where appropriate:
dataset descriptions;
ownership and stewardship;
data definitions;
source systems;
collection methods;
update frequency;
quality characteristics;
known limitations;
transformation processes;
lineage and provenance;
security classification; and
permitted uses.
This documentation should be maintained as part of normal data operations rather than created only for individual AI projects.
10.7 Developing Data Lineage and Provenance
Data lineage should be treated as an important organisational capability because AI outputs may depend on multiple upstream sources and transformations.
If a model produces unexpected results, organisations need to determine whether the problem originated in the source data, a transformation process, an integration step or the AI system itself. Without lineage, identifying the source of a problem can become difficult and time-consuming.
Qlik (2026) identifies lineage as an important mechanism for supporting trustworthy data, while Heil et al. (2021) emphasise documentation of data processes in the context of reproducible machine learning.
Organisations should therefore seek to establish traceable relationships between:
Source data → transformed data → AI-ready representations → models/applications → outputs
This does not necessarily require every organisation to implement the same technical architecture. Rather, it requires sufficient traceability to understand how data contribute to AI outcomes.
10.8 Continuous Data Quality Management
AI data readiness should also change the way organisations approach data quality. Instead of treating quality assessment as a one-time cleansing exercise, organisations should establish continuous monitoring mechanisms.
Priestley et al. (2023) demonstrate that dataset quality is a multidimensional lifecycle concern. This supports an organisational approach in which data quality is continuously evaluated against the requirements of the intended AI application.
Monitoring may include:
completeness checks;
consistency checks;
anomaly detection;
duplicate identification;
validation rules;
changes in data distributions;
missing-data monitoring;
source-system changes; and
investigation of unexpected data patterns.
The objective should not be to eliminate every imperfection. Rather, organisations should determine whether observed data-quality issues are significant enough to affect the intended AI application.
10.9 Monitoring Representativeness and Bias
Organisations should also monitor diversity and representation rather than assessing them only during initial model development.
Bernhardt, Jones and Glocker (2022) demonstrate that dataset bias can arise through multiple mechanisms and that identifying its causes can be complex. This suggests that organisations should avoid treating diversity as a simple numerical measure.
Instead, assessments should consider whether the data adequately represent the population and circumstances relevant to the AI application's intended use.
Questions may include:
Which groups or conditions are represented?
Which may be underrepresented?
How were the data collected?
Could sampling processes introduce systematic limitations?
Are particular groups disproportionately affected by missing information?
Have changes in the deployment environment altered the representativeness of the data?
This approach recognises that diversity is contextual and that representativeness may change over time.
10.10 Maintaining Timeliness in Changing Environments
The principle of timeliness creates an additional organisational requirement: data pipelines must be capable of responding to change.
A dataset that was appropriate when an AI system was developed may become less suitable as customer behaviour, organisational processes or other environmental conditions change. Consequently, organisations should establish monitoring mechanisms capable of identifying when information is becoming outdated or when the relationship between historical data and the current environment has changed.
The appropriate update frequency will depend on the application. Some systems may require near-real-time information, whereas others may appropriately rely on historical data.
The key organisational question is therefore not simply:
How recent are the data?
but:
How recent do the data need to be for this particular AI application?
This distinction reinforces the contextual nature of AI readiness.
10.11 Embedding Security and Privacy by Design
Security should be incorporated into AI development from the beginning rather than added after technical development has been completed.
Janssen et al. (2020) emphasise the importance of governance for trustworthy AI, supporting the integration of security, access and responsible data use into organisational processes.
Organisations should consider security throughout the movement of data from source systems into AI environments. This is particularly important for GenAI and retrieval-augmented generation (RAG), where information may be transformed, indexed and made accessible through new interfaces.
Security requirements should therefore cover:
data classification;
access control;
authentication and authorisation;
protection during storage and transfer;
monitoring of access;
management of sensitive information;
appropriate restrictions within AI retrieval systems; and
ongoing review of permissions.
The principle should be controlled accessibility, allowing legitimate AI use while maintaining appropriate protection.
10.12 Building Discoverable Data Environments
Organisations should also move away from fragmented data environments in which users depend on informal knowledge to locate relevant information.
Searchable catalogues, business glossaries, metadata and provenance mechanisms can provide a more systematic approach to discovery (Qlik, 2026).
This is particularly valuable as organisations expand AI adoption. The number of potential data sources available to AI teams can become too large for individuals to understand through personal experience alone.
A discoverable data environment can reduce duplicated data acquisition, support reuse and make it easier for teams to identify appropriate information before developing new AI pipelines.
However, discoverability should remain integrated with security. Information about sensitive datasets can be made visible through controlled metadata without necessarily granting unrestricted access to the underlying content.
10.13 Developing Data Consumability as a Reusable Capability
Organisations should avoid rebuilding data transformation processes separately for every AI application where common requirements exist.
Instead, reusable pipelines and representations can support multiple AI use cases. However, standardisation should not remove application-specific requirements. Priestley et al. (2023) demonstrate that dataset quality must be considered in relation to the intended machine learning context, while Qlik (2026) emphasises the different requirements of ML and GenAI.
For conventional ML, reusable capabilities may include structured data preparation, feature engineering and validation.
For GenAI, they may include:
document extraction;
segmentation;
semantic representation;
metadata enrichment;
indexing;
retrieval preparation; and
validation of transformed information.
These processes should maintain links to the original information so that transformed representations remain traceable.
10.14 Preparing for Generative AI and Large Language Models
The increasing use of GenAI and LLM applications creates a particularly strong need for organisational data readiness.
Unlike some traditional analytical systems, GenAI applications may need to draw upon broad collections of organisational knowledge, including documents and other unstructured information. Qlik (2026) identifies RAG as an approach for connecting LLMs with organisational information.
This means that organisations should not view the acquisition of an LLM capability as equivalent to having an AI-ready information environment. The effectiveness of such applications depends substantially on the quality, relevance, security, discoverability and consumability of the information made available to them.
Organisations should therefore assess the underlying knowledge environment before expanding GenAI deployment.
Particular attention should be given to:
the currency of organisational knowledge;
document provenance;
access restrictions;
metadata quality;
semantic organisation;
retrieval suitability;
transformation quality; and
mechanisms for updating information.
This reinforces the central argument of the paper that AI capability depends heavily on the quality of the underlying data foundation.
10.15 Developing Cross-Functional AI Governance
AI data readiness cannot be managed effectively by one organisational function alone.
Data engineers may understand technical pipelines but lack detailed knowledge of business context. Domain experts may understand the meaning of information but not its technical transformation. Security specialists may identify risks that are not visible to data scientists, while governance specialists may understand accountability requirements that are not captured within technical workflows.
Janssen et al. (2020) support the need for organisational arrangements that connect data management and trustworthy AI. Organisations should therefore establish mechanisms for cross-functional collaboration.
This may include governance committees, AI review processes, data stewardship networks or project-level readiness assessments. The specific organisational structure will vary, but the underlying principle is consistent: AI data decisions should combine technical, contextual and governance expertise.
10.16 Developing an AI Data Readiness Operating Model
The six principles can be translated into an organisational operating model consisting of five broad activities:
1. Identify
Determine what data are required for the AI use case and identify relevant sources.
2. Assess
Evaluate the data against diversity, timeliness, accuracy, security, discoverability and consumability requirements.
3. Remediate
Address identified weaknesses through data improvement, additional collection, transformation, documentation, governance or architectural changes.
4. Approve
Determine whether the data satisfy the minimum requirements for the intended application.
5. Monitor
Continuously evaluate whether readiness is being maintained as data, systems and operating environments change.
This model transforms the six principles from a conceptual framework into a repeatable organisational process.
10.17 Measuring Organisational Maturity
Organisations should distinguish between the readiness of an individual dataset and the maturity of their overall data capability.
An organisation may have several highly mature datasets while still lacking enterprise-wide processes for maintaining AI-ready data. Conversely, an organisation may have strong governance processes but insufficient technical infrastructure for AI consumability.
Maturity assessment should therefore examine organisational capabilities such as:
governance structures;
data stewardship;
metadata management;
quality monitoring;
lineage;
security;
data engineering;
AI-specific assessment processes; and
continuous improvement.
Qlik's (2026) AI Trust Score may provide one mechanism for communicating progress, but maturity should not be reduced to a single measure. Organisations should understand both where weaknesses exist and why they exist.
10.18 Economic and Operational Implications
Investment in AI data readiness may initially appear to increase the cost and complexity of AI projects. Data documentation, governance, quality assessment and transformation require time and resources.
However, inadequate data foundations can generate significant downstream costs. Poor-quality data may lead to unreliable models, repeated development efforts, additional remediation and loss of confidence in AI systems.
A more mature approach can therefore shift effort from reactive problem-solving towards proactive data management.
The objective is not to make every dataset perfect before any AI experimentation occurs. Rather, organisations should apply controls proportionate to the importance and risk of the intended application.
This supports a risk-based investment model in which greater organisational resources are directed towards data supporting high-impact or high-risk AI applications.
10.19 Organisational Change and Skills
AI data readiness also has implications for organisational skills. The transition towards AI-enabled operations requires employees who understand not only AI technologies but also data quality, governance, security and contextual interpretation.
Technical teams need greater awareness of governance and responsible data use, while governance and business teams increasingly need sufficient technical understanding to evaluate AI data requirements.
The multidisciplinary nature of AI data readiness therefore suggests that organisations should develop shared data literacy across relevant functions.
This does not require every employee to become a data scientist. Rather, it requires individuals involved in AI decision-making to understand the implications of data quality, representation, security, documentation and technical usability.
10.20 Avoiding the "AI First, Data Later" Approach
A central organisational implication of this framework is the need to avoid an AI-first, data-later approach.
Organisations may be attracted by the capabilities of increasingly sophisticated models and assume that technical advances will compensate for weaknesses in their data environment. The literature examined in this paper suggests otherwise.
Priestley et al. (2023) demonstrate the importance of dataset quality to machine learning, while Janssen et al. (2020) position data governance as a foundation for trustworthy AI. These perspectives indicate that sophisticated AI architectures cannot eliminate fundamental problems in the data on which they depend.
AI initiatives should therefore begin with questions such as:
What information is required?
Is it representative?
Is it sufficiently current?
Is it accurate enough?
Is it appropriately protected?
Can it be found and understood?
Can it be consumed by the intended AI architecture?
Only after these questions have been addressed should organisations determine whether a particular AI technology is appropriate.
10.21 Organisational Implementation Roadmap
The framework can be translated into a staged organisational roadmap.
Stage 1: Establish the baseline
Organisations should inventory important data assets, identify existing governance structures and assess current capabilities against the six principles.
Stage 2: Address critical deficiencies
Priority should be given to deficiencies that create significant security, privacy, quality or representational risks.
Stage 3: Develop enabling infrastructure
Organisations should strengthen metadata, catalogues, lineage, quality monitoring, security controls and reusable data pipelines.
Stage 4: Integrate readiness into AI governance
Readiness assessment should become part of the AI project lifecycle rather than a separate initiative.
Stage 5: Establish continuous monitoring
Organisations should monitor changes in data, AI applications and operating environments and reassess readiness accordingly.
Stage 6: Develop organisational maturity
Over time, organisations should move from project-specific remediation towards reusable enterprise capabilities for maintaining AI-ready data.
This roadmap reflects the central argument that AI readiness is a continuous organisational capability.
10.22 Implications for Responsible Innovation
AI data readiness also contributes to responsible innovation. By requiring organisations to examine representation, quality, privacy, security, provenance and usability before deployment, the framework encourages organisations to consider the broader consequences of AI development.
Bernhardt, Jones and Glocker (2022) demonstrate the complexity of dataset bias, while Janssen et al. (2020) emphasise governance as a foundation for trustworthy AI. Together, these perspectives support a model in which responsible AI begins before model deployment and includes decisions about which data are collected, selected and transformed.
Responsible innovation therefore requires organisations to ask not only:
Can we build this AI system?
but also:
Should we use these data for this purpose, and can we govern their use responsibly?
This distinction is fundamental to the concept of AI readiness.
10.23 Limitations and Organisational Challenges
Although the six-principle framework provides a useful basis for organisational implementation, several challenges remain.
First, organisations may struggle to define measurable thresholds for each dimension. The appropriate level of diversity, timeliness or accuracy will vary by application.
Second, some dimensions involve qualitative judgements. Representativeness and contextual relevance, for example, cannot always be reduced to simple numerical indicators.
Third, organisations may face resource constraints. Establishing comprehensive metadata, lineage and continuous monitoring can require substantial investment.
Fourth, governance processes can become overly bureaucratic if they are not proportionate to the risks associated with individual applications.
These limitations reinforce the importance of adopting a risk-based and context-sensitive approach rather than attempting to impose identical readiness requirements on every AI use case.
10.24 Strategic Implications
At a strategic level, AI data readiness should be regarded as part of organisational AI capability.
Organisations that invest in high-quality data foundations may be better positioned to reuse information across multiple AI applications rather than rebuilding data environments for each project. Discoverable and well-governed data can reduce duplication, while reusable transformation and consumability capabilities can accelerate AI development.
The strategic value therefore extends beyond individual model performance. AI-ready data can become an organisational asset that supports experimentation, deployment and adaptation across a portfolio of AI applications.
This reinforces the proposition that the transition to AI is not simply a transition to new models. It is also a transition towards more mature organisational approaches to data.
10.25 Conclusion
The implications for organisations extend across governance, technology, processes, skills and organisational culture. The six principles proposed by Qlik (2026)—diversity, timeliness, accuracy, security, discoverability and consumability—provide a practical structure through which organisations can assess and improve their data foundations for AI.
The analysis indicates that organisations should integrate data readiness into the entire AI lifecycle, establish clear ownership and accountability, invest in metadata and lineage, continuously monitor data quality and timeliness, protect sensitive information and develop reusable mechanisms for data transformation and consumption.
The central organisational shift is from treating data preparation as a project-specific technical task towards treating AI data readiness as an enterprise capability. This requires collaboration between data professionals, AI practitioners, governance specialists, security teams and domain experts.
The literature supports this broader perspective. Dataset quality is multidimensional and lifecycle-based (Priestley et al., 2023); data governance is fundamental to trustworthy AI (Janssen et al., 2020); dataset documentation supports reproducibility and responsible machine learning (Heil et al., 2021); and dataset bias requires careful consideration of how data are collected and represented (Bernhardt, Jones and Glocker, 2022).
Consequently, organisations should avoid an AI-first approach in which model selection precedes assessment of the underlying data. Instead, AI initiatives should begin by establishing whether the relevant information is sufficiently diverse, timely, accurate, secure, discoverable and consumable for the intended application.
Ultimately, the organisational objective should not be to create a perfect dataset, nor to achieve a universally defined AI readiness score. It should be to establish a repeatable, risk-based and continuously improving capability for determining whether data are fit for responsible AI use. Such a capability provides the organisational foundation necessary for sustainable AI adoption and strengthens the connection between data management, trustworthy AI and responsible innovation.
11. Conclusion
The increasing adoption of artificial intelligence, machine learning and generative AI has reinforced the importance of establishing reliable and appropriately governed data foundations. AI systems depend not simply on the availability of large volumes of information, but on whether that information is suitable for the intended application, sufficiently reliable and capable of being governed throughout its lifecycle. This paper has examined this challenge through the six principles of AI-ready data proposed by Qlik (2026): diversity, timeliness, accuracy, security, discoverability and consumability.
The analysis demonstrates that each principle addresses a distinct but interconnected aspect of AI data readiness. Diversity concerns the representativeness of data and the risks associated with underrepresentation and dataset bias. Timeliness addresses whether information remains relevant as operating environments change. Accuracy provides the basis for reliable analysis and modelling, while security addresses the protection, integrity and responsible use of data. Discoverability enables organisations and AI systems to locate, understand and trace relevant information, while consumability ensures that data can be appropriately transformed and incorporated into the technical architectures used for ML and GenAI.
A central finding of the paper is that these dimensions should not be treated as independent checklist items. Their interaction is fundamental to AI readiness. Data may be accurate but outdated, diverse but poorly documented, discoverable but inadequately protected, or technically consumable while lacking sufficient contextual information. Such examples demonstrate that strength in one dimension cannot necessarily compensate for a critical weakness in another.
AI readiness should therefore be understood as a multidimensional and context-dependent property. The suitability of data depends on the intended AI application, the characteristics of the operating environment and the risks associated with the use of the data. A dataset that is appropriate for one application may not be appropriate for another. Consequently, organisations should evaluate the six dimensions against specific AI requirements rather than applying universal standards without consideration of context.
The analysis further establishes that AI readiness is dynamic rather than static. Data sources, organisational processes, user requirements and operating environments change over time. As a result, a dataset that is suitable at the point of initial development may become less appropriate following changes in its environment. This supports a lifecycle approach in which organisations assess, prepare, deploy, monitor, remediate and reassess their data continuously. Such an approach is consistent with the literature emphasising the lifecycle nature of dataset quality and the importance of documenting data collection, cleaning and curation processes (Priestley et al., 2023; Heil et al., 2021).
The paper also identifies an important distinction between AI data readiness and trustworthy AI. AI-ready data are a necessary foundation for trustworthy AI, but they do not by themselves guarantee trustworthy outcomes. Model design, system implementation, human oversight and organisational governance also contribute to AI system performance and trustworthiness. Nevertheless, weaknesses in the underlying data can undermine these wider safeguards. Data governance is therefore integral to AI readiness, consistent with Janssen et al. (2020), who identify governance as a foundation for organising data for trustworthy AI.
The framework also has practical implications for organisations. Rather than treating data preparation as a project-specific technical task, organisations should develop an enterprise capability for assessing and maintaining AI-ready data. This requires clear ownership and accountability, appropriate metadata and documentation, data lineage, continuous quality monitoring, security controls and reusable mechanisms for data transformation and consumption. It also requires collaboration between data engineers, data scientists, domain experts, security specialists and governance functions.
The proposed readiness assessment provides a structured basis for this organisational approach. However, the analysis cautions against reducing AI readiness to a single numerical indicator. Although composite measures such as Qlik's (2026) AI Trust Score may be useful for communicating overall progress, a high aggregate score could conceal critical weaknesses. Organisations should therefore combine dimension-specific assessment with minimum thresholds for critical requirements and broader measures of organisational maturity.
The principal contribution of this paper is consequently to position AI data readiness between traditional data quality and broader trustworthy AI governance. Traditional data quality remains essential, but AI requires additional consideration of representation, temporal relevance, security, discoverability and technical usability. The six-principle framework brings these dimensions together into an integrated data-centred perspective on AI readiness.
Based on the analysis, AI data readiness can be defined as the multidimensional and continuously maintained capacity of data to support a specified AI application by being sufficiently diverse, timely, accurate, secure, discoverable and consumable within its organisational and operational context. This definition emphasises that readiness is not an inherent or permanent property of a dataset. Rather, it is a relationship between data, technology, purpose, context and governance that must be maintained over time.
Future research should therefore move beyond conceptualisation towards empirical validation of the framework. Further studies could investigate how the six dimensions can be measured consistently, examine relationships between data readiness and AI system performance, establish application-specific readiness thresholds and explore how organisations can integrate readiness assessment into large-scale AI governance and data management practices.
Ultimately, sustainable AI adoption depends on more than access to increasingly sophisticated models. It requires organisations to establish data foundations that are reliable, representative, current, protected, understandable and technically usable. AI readiness should therefore be regarded not as a one-time assessment or certification, but as an ongoing organisational capability that connects data quality, governance, technical architecture and responsible AI practice.
References
Bernhardt, M., Jones, C. and Glocker, B. (2022) ‘Potential sources of dataset bias complicate investigation of underdiagnosis by machine learning algorithms’, Nature Medicine, 28, pp. 1157–1158.
Priestley, M., O’Donnell, F. and Simperl, E. (2023) ‘A survey of data quality requirements that matter in ML development pipelines’, ACM Journal of Data and Information Quality, 15(2), pp. 1–39.
Heil, B.J. et al. (2021) ‘Reproducibility standards for machine learning in the life sciences’, Nature Methods, 18, pp. 1132–1135.
Janssen, M., Brous, P., Estevez, E., Barbosa, L.S. and Janowski, T. (2020) ‘Data governance: Organizing data for trustworthy Artificial Intelligence’, Government Information Quarterly, 37(3), 101493. doi: 10.1016/j.giq.2020.101493.
Nature Computational Science (2021) ‘Moving towards reproducible machine learning’, Nature Computational Science, 1, pp. 629–630.
Qlik (2026) The Six Principles of AI-Ready Data: Establishing a Trusted Data Foundation for AI. Qlik Whitepaper.
Contact
Reach out via email for inquiries.
Subscribe to newsletter
info@grcadvisory.ch
© 2025. All rights reserved.