Synthetic data, synthetic trust: navigating data challenges in the digital revolution

合成数据,合成信任:应对数字革命中的数据挑战

阅读:1

Abstract

In the evolving landscape of artificial intelligence (AI), the assumption that more data lead to better models has driven unchecked reliance on synthetic data to augment training datasets. Although synthetic data address crucial shortages of real-world training data, their overuse might propagate biases, accelerate model degradation, and compromise generalisability across populations. A concerning consequence of the rapid adoption of synthetic data in medical AI is the emergence of synthetic trust-an unwarranted confidence in models trained on artificially generated datasets that fail to preserve clinical validity or demographic realities. In this Viewpoint, we advocate for caution in using synthetic data to train clinical algorithms. We propose actionable safeguards for synthetic medical AI, including standards for training data, fragility testing during development, and deployment disclosures for synthetic origins to ensure end-to-end accountability. These safeguards uphold data integrity and fairness in clinical applications using synthetic data, offering new standards for responsible and equitable use of synthetic data in health care.

特别声明

1、本页面内容包含部分的内容是基于公开信息的合理引用;引用内容仅为补充信息,不代表本站立场。

2、若认为本页面引用内容涉及侵权,请及时与本站联系,我们将第一时间处理。

3、其他媒体/个人如需使用本页面原创内容,需注明“来源:[生知库]”并获得授权;使用引用内容的,需自行联系原作者获得许可。

4、投稿及合作请联系:info@biocloudy.com。