Archived transaction data

Turn your archive into a moat, not a liability.

Decades of transaction history can sit in formats that modern analytics and AI workflows cannot meaningfully query.

Aerial view of a Swiss lake and mountain valley

The central risk

An AI strategy cannot learn from history it cannot understand.

An archive can satisfy retention while remaining useless for investigation and analysis. The missing layer is often not more storage. It is a usable account of schemas, relationships, time, lineage and business meaning.

Why this matters now

Historical transactions still have work to do.

The value appears when the organisation needs to investigate a pattern, explain a decision or see a relationship across time—not when the data is merely retained.

Fraud and AML investigation

Long-range patterns may matter, but only when analysts can query transactions with the context needed to interpret them.

Pattern and context

Credit and client history

Risk, value and attrition questions often depend on relationships that span products, platforms and changes in schema.

History across systems

Regulatory and audit response

A retained record is not automatically an explainable record. Retrieval, lineage and interpretation remain separate responsibilities.

Evidence on demand

The false default

The extract-and-hope migration creates a better-shaped archive, not a more useful one.

A one-time export can flatten temporal relationships, freeze one interpretation into a rigid schema and optimise for the first use case anyone thought to ask. The next question starts another extraction project.

The working path

Discover. Model. Govern. Validate. Hand over.

A schema-on-read lakehouse is one possible architecture. The durable principle is to separate retained source evidence from evolving analytical interpretations.

  1. Archive discovery

    Map the evidence

    Identify source systems, formats, periods, joins, retention constraints and the people who can explain what the fields meant at different times.

  2. Semantic modelling

    Keep interpretation explicit

    Represent changing schemas, identifiers and business relationships without pretending the archive was always one clean model.

  3. Load and governance

    Separate source from interpretation

    Place retained data and analytical models under access, ownership and lifecycle controls appropriate to the estate.

  4. AI-readiness validation

    Test useful questions

    Use representative queries to expose missing context, ambiguous joins and model assumptions before a production use case depends on them.

  5. Query handoff

    Leave a usable foundation

    Give authorised teams a documented way to ask new questions without returning to the legacy platform for every answer.

Technology landscape

Choose the architecture after understanding the archive.

Snowflake, Databricks and open table formats can support this pattern. Their suitability depends on the source estate, governance model, access needs and operating ownership.

Analytical platform context

  • Snowflake
  • Databricks
  • Delta Lake
  • Apache Iceberg
  • Object storage
  • Schema-on-read modelling

Source estate context

  • Mainframe records
  • IBM i / AS/400
  • COBOL
  • RPG
  • Relational databases
  • Decommissioned application exports

Decisions, not theatre

Keep the archive distinct from the reporting warehouse.

The archive can preserve source history and support new interpretations while the reporting warehouse continues to serve controlled operational reporting.

What remains source evidence?

Define which retained records must remain distinguishable from later cleaning, enrichment and interpretation.

What can evolve?

Let analytical schemas change as new questions emerge without rewriting the evidential history underneath them.

Who may query what?

Make access, purpose, ownership and retention part of the architecture rather than an afterthought.

What proves readiness?

Use real investigative and analytical questions to test whether the archive carries enough context to support responsible use.

Questions worth asking

Before the programme writes its own answer.

What counts as archived transaction data?

Historical transaction records and their surrounding context: parties, accounts, products, events, statuses, corrections, source identifiers and the schema changes that shaped them over time.

Why not load everything into the reporting warehouse?

A reporting warehouse usually serves defined operational models. An archive has a different job: preserve historical evidence and support questions that may not yet have a stable reporting definition.

Does schema-on-read remove the need for modelling?

No. It delays irreversible modelling choices, but useful queries still require explicit relationships, definitions, ownership and tests.

When is an archive AI-ready?

Not when a platform has loaded the files. Readiness begins when representative questions can be answered with understood data, visible assumptions, appropriate access and accountable review.

A useful first conversation

Test whether the archive is retained, retrievable or genuinely ready to answer a new question.

Bring the system, history and decision the organisation cannot afford to misunderstand. The first step is to establish what evidence exists and what must become reviewable.