In the era of information and the ascent of deep learning, data-driven artificial intelligence (AI) models have progressively transitioned from theoretical concepts to practical applications, demonstrating remarkable capabilities and potential across various sectors. The financial industry is no exception to the influence of AI technology. Numerous AI models have become standard tools in finance, with models based on deep learning starting to make an impact[1], showing promise to surpass traditional models. However, training these deep learning models demands substantial data to fine-tune the model parameters. The advancement of deep learning models in finance is greatly constrained by the issue of data scarcity. In this article, we will analyze the data scarcity challenges that financial AI models face, discuss prevalent data generation techniques, and explore the corresponding opportunities and prospects.
Main Reasons and Impacts for Data Scarcity of Financial Data
- Financial data is limited by its time series nature. The volume of data in financial assets depends on their historical length. For instance, data for individual stocks can only be traced back to their initial public offering date. Currently, the two most advanced domains for deep learning models – language and images – demand training data on the scale of hundreds of millions (see Figure 1). In contrast, the scarcity of commonly used financial datasets is evident, such as stock prices and GDP figures, which contain only hundreds to tens of thousands of data points.

- Privacy and security concerns are also important factors[1,2,4]. Financial data frequently includes a significant amount of sensitive information, such as personal assets and transaction records. To safeguard the security of this information, relevant regulations and policies impose strict limitations on the collection and use of financial data.
- There is also the issue of imbalanced categories in financial data. Financial data can be categorized into various types, and certain categories of data make up a very small proportion in comparison. For instance, less than 0.1% of credit card transaction data may be considered an anomaly[4].
The interplay of these significant factors prevents AI models from fully realizing their potential in financial scenarios. When data is limited, deep learning models tend to overfit the training data rather than capture the main distributions and common features. This leads to suboptimal performance and a lack of generalization ability on new, unseen data. Additionally, when the original data volume is scarce, the low proportion of various categories of data further complicates the model’s ability to accurately learn the characteristics of these categories.
Augmenting Financial Datasets by Data Generation Techniques
In order to alleviate the problem of data scarcity in the financial field, we can use the following common data generation methods to augment financial datasets.
- Data Resampling: This method allows for changing the time frequency of the data, such as converting daily data into weekly or monthly data. Resampling can uncover trends and patterns across different time scales, which proves particularly valuable for analyzing long-term investment strategies.
- Data Transformation: This method generates new data through various transformation operations, such as noise injection, scaling, and time rotation. Data transformation enables the creation of several subtly different new samples from a historical sample, offering enhanced testing opportunities for short-term and time-specific investment strategies.
- Financial Statistical Methods: In the financial sector, numerous classical statistical methods are available to extract specific features from historical data to simulate new data. These methods include the Autoregressive Moving Average (ARMA) and GARCH models, which are designed to capture key characteristics of financial time series data, such as seasonality, trends, and volatility. These models can simulate potential future movements based on historical data, which can be utilized for risk assessment or analyzing market behavior.
- Deep Generative Models: Using deep generative models, such as Generative Adversarial Networks (GANs) and Diffusion models, we can directly learn the distribution of existing datasets. Once the learning process is completed, these models are capable of generating new data samples that closely resemble those in the actual dataset. This process is akin to drawing a sample from a well-known mathematical distribution, such as a Gaussian distribution. Consequently, we can produce seemingly authentic financial market data that can be utilized to simulate market conditions or test financial strategies.
Data resampling offers a straightforward approach to analyze long-term trends without enhancing the diversity of the data. Financial statistical models and data transformations are more sophisticated in terms of simulation and risk assessment, making them well-suited for examining specific characteristics of time series. Deep generative models hold significant potential in the financial sector, as they can generate new data that adheres to the distribution of historical data without merely replicating past events. These models have shown considerable progress in augmenting datasets, thereby enhancing the performance of deep learning models[1,2,3]. Financial enterprises and institutions can adaptively utilize these methods to supply more training data for their AI models, thereby enhancing their accuracy and generalization capabilities across various scenarios.
New Opportunities for AI-Powered FinTech with Massive Data
The vast amount of generated data offers a broader “learning ground” for deep learning, allowing models to extract more complex rules from richer data to achieve more sophisticated functionalities. In the following sections, we will explore the application prospects of generating financial data in the financial sector and its impact on revenue improvement by examining two practical scenarios.
- Market Prediction: Massive and sufficiently diverse generated data can be used to train more accurate forecasting models and identify new market patterns. For instance, generated data can simulate fluctuations in asset prices under specific economic conditions or trading behavior under varying market circumstances. With this generated data, financial companies can train or validate their financial models in a wider range of scenarios, further refining investment strategies to boost returns. According to research by J.P. Morgan, the accuracy of models that predict stock movements can be enhanced by nearly 28% with the aid of generative models[5], leading to a significant increase in rate of returns.
- Risk Management: Using large amounts of generated data for risk management can simulate market conditions more comprehensively than traditional methods. We can create extensive amounts of “synthetic” data that do not pose privacy issues and can be used to simulate various scenarios in the financial market, similar to the “wind tunnel” experiments used for precision instruments. This provides a safe and controllable environment for decision-makers to evaluate and verify strategies before implementation. For example, banks can assess the resilience and reliability of their trading systems under various stress scenarios to mitigate systemic risk, protect customer assets, and comply with regulatory requirements, thereby reducing potential financial losses. The Findex platform has introduced an innovative data-sharing service designed to enhance fraud identification algorithms by providing high-quality generated data. Similarly, J.P. Morgan utilizes generated data to improve the training of its anti-money laundering and anomalous transaction detection systems.
Big data-driven AI models are profoundly impacting various fields, gradually replacing traditional methods and fostering waves of industrial innovation. The use of generated data in the financial industry represents a new “blue ocean.” Supported by massive data volumes, the deployment of AI models in finance is poised to become a reality. According to a 2018 report by Barclays (refer to Figure 2), companies that leverage AI already constitute a significant share of many financial sectors, with approximately 20% of companies involving AI in 80% to 100% of their decision-making processes[6]. With the aid of more generated data, these companies are expected to see substantial improvements in their investment performance.

Future Prospects and Challenges
The Economist emphasizes that “the most valuable resource of the future is not oil, but data.” In the era of big data and AI models, massive financial data is to financial models what massive text data is to large language models like ChatGPT. This will undoubtedly become a top priority for the comprehensive development and implementation of AI models. In recent years, there has been a surge in companies specializing in data generation. As illustrated in Figure 3, the data generation market for North America alone exceeded $200 million as of 2023 and is projected to approach nearly $2 billion by 2030[9].

“Generated data can be shared across companies, departments, and research units to create synergies and enable testing of extreme scenarios not covered by real-world data,” envisions the Alan Turing Institute. According to the FCA Synthetic Data Expert Group, “generating data enhances the ability to protect consumers and encourages beneficial innovation in financial services.” Tucker Balch, President of J.P. Morgan AI Research, has outlined the three key directions and challenges for future research: authenticity, safety, and efficacy.
Reference:
[1] Assefa, Samuel A., Danial Dervovic, Mahmoud Mahfouz, Robert E. Tillman, Prashant Reddy, and Manuela Veloso. “Generating synthetic data in finance: opportunities, challenges and pitfalls.” In Proceedings of the First ACM International Conference on AI in Finance, 2020.
[2] Dogariu, Mihai, Liviu-Daniel Ştefan, Bogdan Andrei Boteanu, Claudiu Lamba, Bomi Kim, and Bogdan Ionescu. “Generation of realistic synthetic financial time-series.” ACM Transactions on Multimedia Computing, Communications, and Applications, 2022.
[3] Wiese, Magnus, Robert Knobloch, Ralf Korn, and Peter Kretschmer. “Quant GANs: deep generation of financial time series.” Quantitative Finance, 2020.
[4] Sethia, Akhil, Raj Patel, and Purva Raut. “Data augmentation using generative models for credit card fraud detection.” In 2018 4th International Conference on Computing Communication and Automation, 2018.
[5] El-Laham, Yousef, and Svitlana Vyetrenko. “StyleTime: Style Transfer for Synthetic Time Series Generation.” In Proceedings of the Third ACM International Conference on AI in Finance, 2022.
[6] Barclays. “Survey: Majority of Hedge Fund Pros Use AI/Machine Learning in Investment Strategies.” 2018.
[7] “Trends in training dataset sizes.” Accessed on 1 January.
[8] “The world’s most valuable resource is no longer oil, but data.” The Economist, May 6th, 2017. Accessed on 1 January.
[9] Grand View Research. “Synthetic Data Generation Market Size, Share & Trends Analysis Report By Data Type, By Modeling Type, By Offering, By Application, By End-use, By Region, And Segment Forecasts, 2023 – 2030” 2023.
The work described in this article was supported by InnoHK initiative, The Government of the HKSAR, and Laboratory for AI-Powered Financial Technologies (AIFT).
(AIFT strives but cannot guarantee the accuracy and reliability of the content, and will not be responsible for any loss or damage caused by any inaccuracy or omission.)