Legal Issues for Data Professionals: In AI, Data Itself Is the Supply Chain
Data is the supply chain for AI. For generative AI, even in fine-tuned, company-specific large language models, the data that is input into training data comes from a host of different sources. If the data from any given source is unreliable, then the training data will be deficient and the LLM output will be untrustworthy. In this sense, the data supply chain is analogous to a manufacturing supply chain. For example, defects or impurities in raw materials or component parts will cause the finished product to fail or not perform optimally for the end consumer. The remedial efforts necessary to identify the root cause of a defect and retrofit a manufactured product are time-consuming and expensive, equally so with data-driven AI products. This article briefly outlines how the specific business goals of the company should inform the legal strategy for maximizing reliability and safe usage of data in each stage of the data supply chain.

