Introduce a practical framework that evaluates licensed search access, page fetching, extraction, model training, and generated output independently instead of collapsing them into a single yes-or-no copyright question.
Imagine a developer in Lisbon on a Tuesday afternoon. She is about to ship an agent that searches the web, reads product pages, extracts a few fields, and writes a comparison.
Legal asks, “Can we use this?”
The release stalls because everyone is trying to compress five separate questions into one yes or no answer. Approval is uncertain, and the launch gets pulled from that week’s release while the team untangles what “this” means.
A more useful review would separate the system into five layers:
1. Search access
Where do the search results come from? Is the search access licensed? What do the provider’s terms allow you to do with the returned links, snippets, and metadata?
This is a procurement and contract question before it becomes a scraping question.
2. Page fetching
Can your system retrieve the page under the site’s terms and applicable law? Does it respect access controls? Are you fetching public pages, authenticated pages, or content behind a paywall?
A licensed search result does not automatically grant permission to fetch every page it points to.
3. Extraction
What are you taking from the page?
Extracting a price, date, or JSON-LD field raises different issues from copying an article’s full text into your own database. Review the amount, purpose, retention period, and whether the extracted material gets redistributed.
4. Model training
Will the retrieved material be used only as temporary context for a response, or stored in a dataset for training or fine-tuning?
Those are different uses and should get separate decisions. This remains a live issue even when a provider adds protections around generated output. The recent MPA and ByteDance deal is a useful example: output-layer protections do not resolve the underlying training-data question.
5. Generated output
What can the user receive?
A system might fetch material legitimately and still produce an output that reproduces too much of it. Citation, quotation limits, similarity checks, and refusal rules belong at this layer.
Back in the hypothetical Lisbon review, the team replaces “Can we use web content?” with a small table: source, permission, fetched material, retained material, training use, and allowed output.
The answer may still be no. But now it is a specific no, attached to one part of the pipeline, instead of a vague no that blocks everything.
This is also how I think about API design for agent tooling. Search, fetch, and extract may sit behind one key, but they remain distinct operations with distinct provenance and risk.
Not legal advice, obviously. Just a framework that makes the conversation concrete enough for counsel, engineering, and product to evaluate the same system.