URL clues from a Scan are not all equal in the software provenance context, but it can be difficult to find the most important URL clues especially in cases like JavaScript where a single file may have a dozen or more URL clues.
The most important example is that a URL from StackOverflow is extremely likely to be relevant for provenance analysis. There should be other patterns where we can identify high value URL clues that we should prioritize/highlight. Perhaps we could assign a score to a URL to indicate how likely it is to represent important origin or license information.
URL clues from a Scan are not all equal in the software provenance context, but it can be difficult to find the most important URL clues especially in cases like JavaScript where a single file may have a dozen or more URL clues.
The most important example is that a URL from StackOverflow is extremely likely to be relevant for provenance analysis. There should be other patterns where we can identify high value URL clues that we should prioritize/highlight. Perhaps we could assign a score to a URL to indicate how likely it is to represent important origin or license information.