Australian Bank Unlocks AI Innovation Without Compromising Customer Privacy

How a Leading Australian Bank Used DataMasque to Enable Secure Data Access for AI Experimentation

One of Australia’s Big Four banks faced a challenge: how do you give data science and AI teams access to realistic customer data without exposing sensitive information? 

The bank’s dedicated AI innovation division was actively evaluating large language models (LLMs) and machine learning tools to improve customer experiences. But every experiment hit the same wall: getting realistic data required months of approvals, and synthetic data generated by AI lacked the real-world nuance needed to validate models with confidence.

DataMasque changed that equation. In a structured Proof of Concept (POC) run within the bank’s AWS sandbox environment, DataMasque demonstrated that sensitive customer data could be de-identified quickly, accurately, and with full referential integrity preserved - giving AI teams the data fidelity they needed without the compliance risk.

The challenge

The bank’s AI innovation lab was built to move fast to prototype, test, and validate ideas before committing to large-scale engineering investment. But a fundamental tension was slowing everything down: almost every meaningful AI use case required customer data, and accessing that data safely was a months-long process.

Teams seeking data for non-production experimentation faced a complex process: secure funding approval, navigate multi-stage governance committees, locate data ownership, assess privacy risk under Australia’s Privacy Act, formally request masking, and wait. By the time data arrived, project windows had often closed and critical relational linkages between datasets had frequently been severed in the masking process, undermining analytical value.

The only faster alternative was to use an LLM to generate synthetic data. But this approach introduced its own failure mode: synthetically generated data looks plausible but lacks the real-world distributions, edge cases, and statistical patterns that give AI models their predictive power. The gap between synthetic data and real customer behaviour was exactly where AI pilots failed to survive contact with production.

The solution

The bank partnered with DataMasque to evaluate whether automated, intelligent data masking could bridge the gap between speed and safety. The evaluation was structured as a formal Proof of Concept (POC), guided by a single central question:

"Can DataMasque safely deliver realistic, nuanced data for AI experimentation without compromising customer privacy or data integrity?”

The POC ran entirely within a sandboxed AWS environmentent. DataMasque’s subject matter experts guided the bank’s Test Data Management (TDM) team through the platform, with the team operating the tool independently from day one.

Five distinct test scenarios were evaluated across the end-to-end data masking workflow:

  • DataMasque platform deployment in the bank’s cloud environment
  • Synthetic data creation that realistically mimics source data
  • Automated sensitive data discovery and classification
  • Sensitive column masking with anonymised, referentially consistent substitutes
  • End-to-end data integrity verification after masking

A core design principle of the evaluation was repeatability. Once a masked dataset was produced, it could serve as a ‘golden copy’ amd be used across multiple AI experiments and model iterations without requiring any additional approvals or re-processing.

Results

All five hypotheses in the Proof of Concept passed.

✓  DataMasque deployed successfully in the bank’s AWS sandbox environment, with no infrastructure or configuration blockers.

✓  Synthetic data generation produced datasets that realistically mimic source data, preserving structure, statistical patterns, and referential integrity.

✓  Automated sensitive data discovery identified and classified PII and sensitive fields across structured datasets with high accuracy without manual column-by-column review.

✓  Masked datasets replaced sensitive fields with anonymised, realistic, and referentially consistent substitutes, making data suitable for non-production and AI use without re-identification risk.

✓  Referential integrity was maintained across all linked tables and systems after masking, which was a critical requirement for AI model training and multi-dataset experiments.

Beyond the five core hypotheses, the evaluation generated a key qualitative insight that has broad implications for how large enterprises approach AI data provisioning.

The ability to produce a single, high-fidelity masked dataset and reuse it across multiple independent AI experiments represents a step-change in how data-intensive teams can operate. Rather than re-requesting data for every new test cycle and absorbing the compliance overhead that entails, teams can provision once and experiment continuously.

Evaluators noted that DataMasque was both easy to repeat and immediately useful to end users.

Looking forward

The success of the POC accelerated the bank’s path to a full production pilot. Building on the AWS sandbox evaluation, the next phase extends DataMasque’s deployment to production staging data within the bank’s on-premises environment.

Further ahead, the bank’s AI innovation team is exploring DataMasque’s applicability to unstructured data, opening up new categories of AI use cases, including document intelligence, natural language processing, and analysis of customer interaction data.

The broader vision is a data provisioning model in which any AI team can access a realistic, privacy-safe dataset within days rather than months,, without sacrificing the data quality that determines whether an AI model succeeds or fails in production.

‍