SEPTEMBER 24, 2026 - 4 MIN READ
Preparing for the AI era in FinServ: Why tokenization is the next big thing in data security
- Data Governance
- Data Management for AI
- Data Security
- Tokenization

Vince GoveasDirector, Product Management, Capital One Software

Derek BaldusDirector, Data Risk and Privacy, PwC US
As the global financial community prepares to gather at Sibos, one topic is increasingly prominent on every executive agenda: scaling enterprise AI securely.
In the rush to deploy generative AI and advanced machine learning models across banking, payments and capital markets, financial data has become an institution's greatest strategic asset, but also one of its largest liabilities. Safely exchanging that information with partners and credit bureaus remains a primary source of vulnerability. Current security policies are often tied to the environment rather than the data itself, meaning once data leaves the internal ecosystem for real-time credit acquisition or collaborative modeling, control can be lost. Institutions face the daunting task of confirming that sensitive identifiers, like Social Security Numbers (SSNs), remain cryptographically locked even when processed by third-party applications.
Traditionally, strategy leaders have faced a crossroads: increase data utility to drive AI performance, or tighten privacy and security controls to mitigate risk. In a sector defined by stringent regulations (NYDFS Part 500, PCI-DSS, FCRA, CCPA) and escalating cybersecurity threats, that compromise is no longer viable. Financial institutions should have a security paradigm that can enhance analytical output while removing the exposure of sensitive underlying data.
Ahead of our speaking session at Sibos on September 28th, we are diving into the key findings from our joint AI study which demonstrates how enterprise data tokenization can provide a clear path forward.
The core finding: 2x the model accuracy
When institutions build AI models for mission-critical financial applications, such as real-time fraud detection, credit risk modeling, AML transaction monitoring or customer churn prediction, data precision matters significantly.
General-purpose AI models are probabilistic—they guess. In mortgage lending or credit scoring, "mostly right" can create buyback risk. Without a data layer providing deterministic, machine-validated signals, AI efficiency remains a compliance liability.
The primary takeaway from our joint study addresses this friction directly for financial data and security leaders: AI/ML models trained on tokenized data were nearly twice as accurate as those using traditional data masking techniques.
This is a competitive leap forward for quantitative analysts, CISOs and enterprise data architects, as it represents the clear line between a high-precision AI model that directly safeguards revenue and an underperforming model hampered by degraded data inputs.
Why traditional masking falls short in financial services
For years, techniques like static or dynamic data masking have served as the standard approach for de-sensitizing data across non-production environments. However, masking can disrupt the referential integrity of complex AI models.
If a masked primary key, account number or transaction identifier is altered arbitrarily across separate ledgers and relational databases, the critical links between data points can break down. Financial AI models depend heavily on these deep relationships, such as connecting an account number across wire transfers, merchant category codes and historical risk scores. When those links are severed, the AI fails to surface subtle patterns, leading to things like false positives in fraud detection and inaccurate risk scoring.
Tokenization: the enterprise advantage for FinServ
Unlike legacy masking, vaultless data tokenization engines like Capital One Software’s Databolt, preserve both format and referential integrity without sacrificing performance at scale. Databolt brings Capital One’s proven, in-house technology to the broader financial market to help enable organizations to scale data operations without compromising control.
By safeguarding sensitive customer and financial data at the source with a patented tokenization engine, institutions can train models and execute analytics quickly using non-sensitive tokens.
Tokenized financial data delivers three vital properties:
Non-sensitive: Tokens carry no raw sensitive value; if intercepted, they hold little to no utility to attackers.
Format preservation: Tokens retain original structural formats (e.g., a 16-digit credit card number or a standard SWIFT IBAN retains its exact format), preventing legacy core banking software and downstream applications from breaking.
Referential integrity: The exact same sensitive value is mapped consistently to the same token enterprise-wide. A customer’s tokenized identity remains unified across transactional logs, risk engines and cloud data warehouses, enabling AI models to detect real-world financial patterns seamlessly.
The mandate for FinServ and data leaders
Modern enterprise security is evolving away from building higher perimeter walls toward intelligent, asset-level de-risking.
For data architecture and cloud ops leaders: Tokenization can simplify cloud migration by reducing audit scope. Replacing raw financial data with format-preserving tokens can reduce compliance complexity and help lower cloud risk exposure. By de-identifying data at the point of ingestion, institutions "shrink the box" for compliance audits, removing systems handling only tokens from the cardholder data environment (CDE) and significantly lowering operational overhead.
For data privacy and compliance teams: Tokenization can serve as an elite data minimization strategy. It can keep raw PII and financial records out of high-risk operational environments, simplifying adherence to evolving global regulations. Furthermore, logging each access request, query and policy decision can create audit trails, enabling compliance with various reporting mandates and requirements.
For data strategy and CISOs: Tokenization can elevate security from a restrictive cost center to a core business enabler that powers faster, safer AI adoption. By maintaining data quality and privacy at the source, tokenization helps establish responsible AI guardrails necessary to satisfy various governance and compliance requirements.
Join us at Sibos
Attending Sibos on Sept 28th? Add our co-presented session to your agenda. We will be presenting at 10am, 11:30am, 2pm and 3pm EST at the PwC Club and Studio (Miami Beach Convention Center, Rooms 201 - 203). Or reach out to us prior and connect directly with our team.
We look forward to further exploring the findings of our joint study and discussing how a tokenized data architecture can help accelerate your enterprise AI initiatives.

Vince Goveas
Director - Product Management, Capital One Software
Vince is a product leader specializing in building secure solutions. He has experience leading AI/data governance, building data protection infrastructure and pioneering tokenization at global tech companies.

Derek Baldus
Director - Data Risk and Privacy, PwC US
Derek is PwC’s chief architect for data tokenization and encryption services. He supports Fortune 100 companies across industries to secure their AI-enabled analytics platforms through tokenization.
Footnotes
DISCLOSURE STATEMENT: © 2026 Capital One. Opinions are those of the individual author. Unless noted otherwise in this post, Capital One is not affiliated with, nor endorsed by, any of the companies mentioned. All trademarks and other intellectual property used or displayed are property of their respective owners.
