Paper deep dive
IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction
Henry Bodwell, Hong Yang, John C. Simeone, Kelvin Gorospe, Bella Sullivan, Lana Huang, Jessica Gephart, Sandy Aylesworth, Molly Masterton, Naren Ramakrishnan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 6/21/2026, 2:13:21 AM
Summary
IUU+DB is a large language model (LLM)-driven system designed to build a global incident database for Illegal, Unreported, and Unregulated (IUU+) fishing, seafood fraud, and labor abuse. The system processes heterogeneous documents (news, academic papers, government reports) through a multi-stage pipeline involving document pre-processing, source classification, and a specialized KDE (Key Data Element) extraction module. It aims to provide actionable insights for regulators, researchers, and industry by identifying geographic hotspots, behavioral patterns, and supply chain risks across more than 140 countries.
Entities (7)
Relation Signals (5)
IUU+DB â extracts â KDE
confidence 100% ¡ The KDE Extraction Module is the workhorse of the IUU+DB system.
IUU+DB â uses â LLM
confidence 100% ¡ IUU+DB, a large language model (LLM)âdriven system...
IUU+DB â utilizes â DSPy
confidence 100% ¡ For our implementation we again utilized an LLM with DSPy.
IUU+ â includes â Seafood Fraud
confidence 90% ¡ IUU+ to capture a broader suite of fisheries sector environmental and associated supply chain trade-related crimes...
IUU+ â includes â Labor Abuse
confidence 90% ¡ IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in areas that lack applicable laws. We propose the term IUU+ to capture a broader suite of fisheries sector environmental and associated supply chain trade-related crimes and behaviors. Although IUU+ activity is widely recognized as a serious threat to marine ecosystems, markets, and livelihoods, a quantitative understanding of these incidents, e.g., their frequency, geography, species, actors, and patterns in the type of illicit activity, remains difficult to obtain. We propose IUU+DB, a large language model driven system for building a global incident database of IUU+ activity. The system ingests heterogeneous documents, classifies whether they describe relevant incidents, extracts key data elements such as actors, locations, species, vessels, violations, and enforcement outcomes, and supports deduplication and trend analysis. Case studies and validation results show that IUU+DB can help organize fragmented evidence, surface geographic and behavioral hotspots, support fisheries-domain specific research in academia and non-government organizations, assist source and species risk assessments for industry, and provide support for policy implementation and targeted enforcement efforts to government agencies.
Tags
Links
- Source: https://arxiv.org/abs/2606.18181v1
- Canonical: https://arxiv.org/abs/2606.18181v1
Trouble viewing inline? Open PDF directly â
Full Text
53,275 characters extracted from source content.
Expand or collapse full text
IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction Henry Bodwell henrybod@vt.edu Virginia Tech United States Hong Yang yanghd@vt.edu Virginia Tech United States John C. Simeone simeoneconsulting@gmail.com Simeone Consulting, LLC United States Kelvin Gorospe kdgorospe@gmail.com Independent United States Bella Sullivan isullivan@nrdc.org Natural Resources Defense Council United States Lana Huang lrhuang@uw.edu University Of Washington United States Jessica Gephart gephart@uw.edu University Of Washington United States Sandy Aylesworth saylesworth@nrdc.org Natural Resources Defense Council United States Molly Masterton mmasterton@nrdc.org Natural Resources Defense Council United States Naren Ramakrishnan naren@vt.edu Virginia Tech United States Abstract Illegal, unreported, and unregulated fishing (IUU) traditionally refers to fishing activities that violate applicable laws or occur in areas that lack applicable laws. We propose the term IUU+ to capture a broader suite of fisheries sector environmental and asso- ciated supply chain trade-related crimes and behaviors. Although IUU+ activity is widely recognized as a serious threat to marine ecosystems, markets, and livelihoods, a quantitative understanding of these incidents, e.g., their frequency, geography, species, actors, and patterns in the type of illicit activity, remains difficult to ob- tain. We propose IUU+DB, a large language model (LLM)âdriven system for building a global incident database of IUU+ activity. The system ingests heterogeneous documents, classifies whether they describe relevant incidents, extracts key data elements such as actors, locations, species, vessels, violations, and enforcement outcomes, and supports deduplication and trend analysis. Case studies and validation results show that IUU+DB can help organize fragmented evidence, surface geographic and behavioral hotspots, Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conferenceâ17, Washington, DC, USA Š 2018 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X support fisheries-domain specific research in academia and non- government organizations, assist source and species risk assess- ments for industry, and provide support for policy implementation and targeted enforcement efforts to government agencies. Keywords Illegal, unreported, and unregulated (IUU) fishing, Large Language Models (LLMs), Information Extraction. ACM Reference Format: Henry Bodwell, Hong Yang, John C. Simeone, Kelvin Gorospe, Bella Sul- livan, Lana Huang, Jessica Gephart, Sandy Aylesworth, Molly Masterton, and Naren Ramakrishnan. 2018. IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM- driven Information Extraction. In . ACM, New York, NY, USA, 10 pages. https://doi.org/X.X 1 Introduction Illegal, unreported and unregulated fishing (IUU) undermines sus- tainable fisheries management which relies on capping harvests, limiting the use of destructive gears, limiting (or minimizing) har- vest of juveniles, and avoiding fishing in sensitive ecosystems, among other regulations. While it is difficult to quantify the full extent of illegal fishing, a 2009 estimate of global losses from IUU fishing amounted to $10 bn and $23.5 bn annually, representing between 11 and 26 million metric tons [2]. While IUU fishing encompasses a broad set of activities, there is a broader set of associated or analogous activities across the seafood sector, including illegal aquaculture, mislabeling, labor abuses, and trade prohibitions and sanctions evasion. We therefore offer a more inclusive term (IUU+) to capture the broader suite of arXiv:2606.18181v1 [cs.IR] 16 Jun 2026 Conferenceâ17, July 2017, Washington, DC, USABodwell et al. human, environmental, and associated supply chain-related crimes and behaviors. IUU+ behaviors threaten both people and nature, and undermine key aspects of the UN Sustainable Development Goals [3]. The consequences of IUU+ behaviors and their associ- ated crimes are far-reaching: destruction of marine and freshwater ecosystems, loss of human livelihoods and socioeconomic security, safety and human rights of fishing communities, loss of local and national revenues, and unfair competition to legitimate fishermen and businesses [37]. And the impacts of IUU+ behaviors extend beyond the fisheries sector as they have been linked to transna- tional organized crime, illicit financial trade such as trade-based money laundering, and circumvention of trade policies such as tariffs, sanctions, and import prohibitions [24]. To deter, detect, and combat IUU+ behaviors and associated trade, governments have responded with a range of regulatory ini- tiatives that span from on-the-water fisheries monitoring, control, and surveillance (MCS) procedures, to port controls (e.g., Port State Measures Agreement, PSMA), to international trade and import control schemes (e.g., U.S. Seafood Import Monitoring Program, and the European Unionâs IUU Fishing Regulation) [10,28,44,45]. In addition, members of the seafood industry partnered with inter- national organizations to develop a framework for interoperable catch and supply chain data standards through the Global Dialogue on Seafood Traceability (GDST) [13]. Many of these programs rely on the collection and transfer of key data elements (KDEs) across supply chain nodes such as vessel identity, harvest location, species, harvest weight, fishing authorization number, and chain-of-custody information [6, 9, 16]. Despite widespread consensus that IUU+ fishing and related crimes occur globally, there does not exist a comprehensive incident database that systematically catalogs incidents across jurisdictions, source types, and categories of illicit behavior. Some information is reported in the news media while others are scattered across government enforcement actions (e.g., arrests, case filings) and still others are in detailed academic publications and NGO reports. The lack of such a comprehensive incident database makes it difficult to obtain a snapshot of fisheries crime, emerging trends and patterns, and where to invest limited surveillance and enforcement resources. To address these needs, we developed IUU+DB, a global incident database for illegal fishing, seafood fraud, labor abuse, and related illicit activities. Our goal is to leverage advances in LLM-driven information extraction technology to capture incidents, their at- tributes, taxonomies of misconduct, and relevant KDEs. Such a resource could help provide near-real-time evidence of IUU+ be- haviors and how these behaviors are expressed, differentiated, and connected across species and geographies. Such a system would support consistent analysis and provide actionable evidence for researchers and policymakers. For example, under the U.S.â im- port control regulation (SIMP), a recent proposal by the overseeing regulatory agency offers the potential to have a two-tier system for seafood import KDE reporting requirements where the agency could periodically review and adjust the species list for each tier based on risk analysis, moving species between tiers as needed [27]. For such a system to work, information and data related to risks of IUU fishing, seafood fraud, and potential labor abuses would need to be tracked in near-real time through a comprehensive incident database such as the one described here [38]. While LLMs have made it easier to rapidly prototype systems such as is presented here, reliable extraction of incidents remains challenging. Source documents can be alternatively scientific or colloquial, considerably lengthy, feature nuanced distinction be- tween IUU+ categories, and contribute a range of potential KDEs. Deduplication (i.e., knowing whether multiple extractions refer to the same incident) is crucial in event reconciliation. While there have been efforts to catalog information on illegal wildlife trade broadly, most work to-date has been focused on enumerating the range of scientific taxonomy involved and is often limited to a spe- cific geographic region. Additionally, these datasets do not contain insights on the underlying crime or illicit behavior committed thus making it difficult to use these datasets to assist in building inter- vention methods [23,40,52]. Database efforts that try to capture multiple dimensions about an incident and the underlying crime are presently manually curated [4,42]. There is a need for a scalable non-manual way to catalog and classify incidents. Our contributions are: (1) We present IUU+DB, an LLM-driven information extraction framework that automatically discovers, extracts, and clas- sifies structured incident data from heterogeneous sources including news articles, academic papers, NGO documents, and government reports. (2) Through IUU+DB, we are now able to impute global trends in illegal fishing, across more than 140 countries. (3)Finally, we show through case studies how IUU+DB can pro- vide actionable insights for regulators, streamlining seafood supply chain risk assessments. 2 Background and Related Work IUU+ fishing and associated trade includes, but is not limited to, activities such as harvesting that violates national or international laws (illegal), fishing that avoids required reporting or monitor- ing (unreported), fishing that occurs in areas or under conditions where effective management and enforcement do not exist (unregu- lated), theft or other illicit fish farm operations (illegal aquaculture), species substitution and mislabeling (fraud), forced labor (labor abuses), and tariff evasion (trade barrier and sanctions evasion). Tracking and quantifying IUU+ activity is a difficult task [43]. No comprehensive way to track IUU+ behaviors and associated trade presently exists. Prior research on tracking IUU+ is scattered and is either specific to IUU+ behavior-type [22,35,41] or is specific to a particular geography or species [14, 53]. Central to how we organize our information classification and extraction is a granular taxonomy of IUU+ behaviors and KDEs. To enumerate IUU+ behaviors, we surveyed the literature and de- veloped a list of 41 behaviors that fall within the seven themes enumerated above. Regarding KDEs, we conducted a literature review on seafood-specific supply chains to identify KDEs that are either already established or are proposed in the literature to fill identified information gaps. We developed 14 groups of KDEs, which, taken together have over 100 KDE fields (see Appendix A). LLM-Driven Information Extraction. The transition from pre- trained language models such as BERT to LLMs, has enabled an increase in the possible complexity of data extraction as the prob- lems move from strictly extractive to generative [54]. However, IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information ExtractionConferenceâ17, July 2017, Washington, DC, USA Figure 1: Framework of IUU+DB. these generative IE methods have notable issues when it comes to reliability and reproducibility [7]. To account for this, earlier work has looked at ways to limit hallucination through schema and grounded ontologies [15,21]. Other approaches extract the exact supporting context alongside the generative IE to evaluate grounding against after the extraction [39], although this can be costly to verify. 3 Methodology Functionally, we set the minimum information threshold for what qualifies as an incident of IUU+ fishing and associated trade to the common data elements used in crime script analysis: who (ac- tors), what (methods), when (time), and where (location). To ensure broad coverage of IUU+ incidents, we designed our information extraction pipeline around five primary sources: (i) Government documents, such as memorandums and court documentation, (i) Non-Governmental Organizational reports and press releases, (i) News articles, (iv) Academic Papers, and (v) Industry journals. These five sources span a variety of input lengths and styles, and we also hypothesized that they will have differing foci of IUU+ incidents. 3.1 Framework Design IUU+DB is an end-to-end pipeline for turning unstructured docu- ments into structured incident records. It ingests documents, classi- fies them by source type, and extracts relevant information into a standard output schema. The pipeline is guided by an input schema that defines the document categories to recognize and the key data elements (KDEs) to extract. Fig. 1 gives an overview of the IUU+DB pipeline. Based on the KDEs, IUU+ types, and behaviors identified by our team outlined in Appendix A, we classify each source into one of three main source types: (i) sources that describe one or multiple incidents, (i) sources that are related to IUU+ fishing but do not pertain to an incident, and (i) sources unrelated to IUU+ fishing. For each incident, the system extracts information from 14 KDE groups covering 100 possible fields, classified by 7 possible IUU+ types, and 41 behaviors. For sources classified as related to IUU+ but that do not describe a specific incident, the system ex- tracts a set of 3 KDEs. To manage the large extraction task possible, the pipeline first aims to ascertain which KDE groups are present in a document and then extracts fields from only those groups. This effectively reduces the schema information the LLM needs to hold in context at any given moment. 3.2 Document Pre-Processing Module Document Pre-Processing Module is the first step of IUU+DB and accepts plain text, URLs, and PDFs. For PDFs we use Pytesseract 1 to perform Optical Character Recognition (OCR) to render documents ingested into plain text. For URLs we first attempt to extract text data with Playwright 2 ; if that fails, we use an LLM to extract the core text. Once the text is extracted, the system removes format- ting artifacts, checks for duplicate documents using a text hash, and skips documents that have already been processed. Finally, IUU+DB splits the cleaned text into smaller chunks, embeds them using Qwen3-VL-Embedding-2B 3 , and stores them in a vector data- base for later semantic search. We chose this Qwen model as it supports a token length of up to 32k which is more than enough and supports input that includes a mixture of text and image data, so in the future we can incorporate embedded figures from sources. This step can be parallelized so that large document collections can be processed efficiently. 3.3 Source Classification Module The Source Classification Module determines what kind of informa- tion a document contains. Given a document and a set of possible source types, the module first identifies passages that appear rele- vant to IUU+ activity. Our framework uses few shot prompting to classify the docu- ment as one of the following: a source describing an IUU+ inci- dent, a source related to IUU+ fishing but with either more general information or not enough sufficient information provided to be considered a specific incident, or a source unrelated to IUU+ fishing. Because a single document may describe more than one incident, this module also extracts and stores the passages associated with each incident. This helps the KDE extraction module process each incident separately and reduces confusion in later stages. 1 https://pypi.org/project/pytesseract/ 2 https://playwright.dev/python/ 3 https://huggingface.co/Qwen/Qwen3-VL-Embedding-2B Conferenceâ17, July 2017, Washington, DC, USABodwell et al. 3.4 KDE Extraction Module The KDE Extraction Module is the workhorse of the IUU+DB sys- tem. It takes the source document, scope classification, as well as optionally relevant passages and a list of KDE groups. The extrac- tor then makesíruns whereíis the number of relevant passages, or 1 if the classifier did not include them. In each run the extrac- tor first determines the presence of each KDE group given before running the extraction of each KDE fields within the group. If the KDE groups were not given the module defaults to extracting ev- ery field. For our implementation we again utilized an LLM with DSPy 4 . DSPy helps us define a structured schema for extraction and encourages logical consistency across fields. For example, if the system identifies a seafood product, we expect it to also identify the species from which that product is made. From the three source classification, we identify two schema to extract: one for incidents, and another albeit much smaller schema for related to IUU+ fishing which tracks species, countries, and companies mentioned. We do not capture any additional information from unrelated articles. In our experience, we found that some KDEs are relatively rare, and unlikely to occur together within the same incident (e.g. It would be uncommon for there to be a single IUU+ incident as we have described them, to include data about both catch AND aquaculture.) We grouped KDEs therefore by those most likely to co occur with one another and extract them in the modules runs. After extract- ing the KDEs we then classify the IUU+ behavior exhibited in the incident. 3.5 Deduplication and Trend Identification After extraction, the incidents are aggregated in two ways, first by merging duplicate incidents and secondly by grouping related incidents into trends. The De-Duplication and Trend Identification modules first identifies similar incidents using cosign similarity against their sources, we then check whether the duplicate can- didates have matching identification information such as vessel name, or id and matching event or enforcement dates. If the two incidents pass those filters and have high cosign similarity, the module merges the extracted KDEs. In the cases of conflicts, the module queries all attached sources for the relevant KDE field in order to get a consensus. 3.6 Prompt Optimization Using DSPyâs prompt optimizers and manually validated data we automate the iterative task of prompt fine tuning. we utilized MIPROv2 [33] to iteratively finetune the prompt and field descrip- tions themselves. For our metric we assigned rewards for having a correctly filled KDE field with high cosine similarity between the extraction and the validated field. To penalize hallucinations we negatively reinforce incorrect extractions. 3.7 Datasets We collected data for this project in two primary ways, directly from APIs and from websites via webscrapers. We used the following APIs: NewsAPI [26], SerpAPI [36], and CORE API [20]. And also tar- geted the following websites via webscrapers: MongaBay [25], US 4 https://dspy.ai/ Figure 2: Web UI of the IUU+DB system. DOJ [51], Oceana [30], US NOAA Fisheries [29], and Undercurrent News [49]. To create queries for candidate source discovery, we used the previously defined IUU+ types and behaviors to build queries that would capture the wide range of behaviors (see Appendix A). For webscraping, we separated websites into two groups by how specific the sites were to the fish-seafood sector. For more gen- eral sites, e,g., MongaBay and DOJ, we searched for the following key words: "illegal fishing", "IUU fishing", "overfishing", "fishing violation", "illegal catch", "seafood fraud", "forced labor fishing", "illegal aquaculture", and "illegal seafood sanctions". In all these searches, our aim was to capture any relevant articles capturing the illicit behavior as well as narrow the scope to only fishing and associated trade. Since the second group of sites focus solely on the fish-seafood sector (e.g., nongovernmental organization (NGO) websites, NOAA Fisheries, and Undercurrent News), it was the case that if the article referred to an infraction it was highly likely to be about IUU+ behaviors. As such, we selected the following keywords: "violence", "investigation", "coercion", "arrest", "charge", "indict", "fined", "enforce". With the three APIs we can use more complex logic and selected the following query: "IUU or seafood mislabeling or ((Transhipment or sanctions or unfree labor or workplace violations or wage theft) and (ship or seafood or fish) and (arrest or criminal investigation or indictment or search or seizure or fine))". From NewsAPI, instead of using keywords, the query was made up of "concepts" which utilize a semantic searching mechanism allowing us to pull articles from different languages, and related articles that may not use the exact key words. 3.8 Web Interface In addition to exposing APIs for general use and further analysis, we developed a web interface to help subject-matter experts analyze, edit, and audit the data. The interface has five main pages. The Insights page shows statistics and figures from the database. The Sources page allows users to read and explore the documents that have been ingested. The Incidents and Related to IUU+ pages allow users to review the KDEs extracted under each schema. The Upload page allows administrative users to add new documents to the database from plain text, URLs, or PDFs. The Sources, Incidents, and Related to IUU+ pages also include version history, so administrators can track changes made to each record (see fig 2). IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information ExtractionConferenceâ17, July 2017, Washington, DC, USA Figure 3: IUU+ Incidents Extracted by Country. Table 1: Incidents extracted via Hotspot Analysis. CountryIncident BrazilMurder of Bruno Pereira & Dom Philips [1] UKIllegal Spear Fishing [34] Alaska, USA North Star Fishing Co. Obstruction [50] Gulf of Mexico, USAIllegal Red Snapper Harvesting [47] Nova ScotiaIllegal Eel-Elving Harvesting [19] British Columbia30k Pounds of Tuna Seized [12] Sri LankaNavy Arrests 14 Indian Fishermen [11] PhilippinesPoaching in Marine Reserve [5] 4 Experimental Results IUU+DB ingested articles over 11 years from 2,472 sources across 143 different countries yielding a total of 8,435 incidents. Our eval- uation is focused on addressing the following questions: â˘Can IUU+DB be used by SMEs to identify global trends in illegal fishing? (§4.1) â˘Does the distribution of KDE extractions follow our expecta- tions of IUU+ behavior types? (§4.2) â˘Can biclustering on (KDE fields, Behavior types) reveal in- sights into reporting patterns? (§4.2) â˘Can IUU+DB correctly identify and extract IUU+ incidents and KDEs? (§4.4) ⢠What are the most surprising incidents tracked? (§4.3) â˘Does IUU+DB improve extraction/classification tasks vs baseline LLMs like GPT-4o-Mini? (§4.4) â˘Does IUU+DB demonstrate improved performance vs. newer/more powerful models such as GPT-5.4? (§4.5) 4.1 Hotspot Analysis We find that the data set and IUU+DB is capturing a wide scope of incidents, from high profile cases [1,48], to regional conflicts [11, 18, 47], to local news [8, 34]. From Fig. 3 we can see that while English speaking nations are highly represented, the dataset has captured events from every continent and nearly every coastal state. The two largest trends we extracted were Illegal Red Snapper harvesting in the Gulf of Mexico and Sri Lankan-Indian fishing rights conflicts. Figure 4: KDE Group Extraction Rate. 4.2 KDE Prevalence Analysis For a given incident, IUU+DB extracts roughly 20% of the fields in the full schema; see Fig. 4. This is not surprising given the size of the schema and the kinds of articles that make up most of the database. We next asked whether IUU+DB fills the fields we would expect for each IUU+ classification. For this purpose, we first analyzed the KDE Presence detection module, which decides whether the system should search for fields within a given KDE group. We then measured KDE group presence across IUU+ types. Figures 4 and 5 show that some KDE groups, such as event and compliance information, appear across many IUU+ types. However, behavior- specific KDEs are much more common in the categories where they are expected. For example, labor standards and crew information increase from 7% and 16% across all cases to 81% and 50% in labor abuse cases. Similarly, aquaculture information increases from 3% across all cases to 70% in aquaculture-related cases. The second analysis we used biclustering to find coordinated occurrences of KDEs and IUU+ sub behaviors. We use the FP-Max algorithm 1 wherein each transaction is modeled as a KDE with one or more IUU+ subtypes. To account for data imbalance, we ran the biclustering in two stages. The first stage was conducted with a minimum support threshold of 0.25 (which yields primarily transactions of IUU+ type of "Illegal Fishing"), and the second with a minimum support threshold of 0.1 after filtering out transactions from the first step. The maximum set bicluster we found was made up of the following 10 KDEs all from the Event group: event and enforcement country, event and enforcement location, and event and enforcement location category, enforcement category, primary offender, event date, and resolution. This cluster was supported by 45% of transactions and by 36 unique IUU+ sub-behaviors. In the second run, there were four clusters each with 10 KDEs, with the most common being the made up of the same 10 KDEs from the first run. The other three were the same as the set above, less event date and with a different KDE instead that is highly relevant to a specific IUU+ behavior or type. Such as, species or product information for Species Fraud and Mislabeling, or crew information information in cases of Labor Abuse. 4.3 Maximum Entropy Analysis We aimed to determine the most surprising illegal fishing incident that occurred within our dataset. For this purpose, we employed 1 https://rasbt.github.io/mlxtend/user_guide/frequent_patterns/fpmax/ Conferenceâ17, July 2017, Washington, DC, USABodwell et al. Figure 5: KDE Group Extraction Rate: Aquaculture. Table 2: Top 6 Results from Maximum Entropy Analysis. CountryQtrIUU+ Sub BehaviorZ Score USA2022Q2Keeping Undersized fish26.6 China2022Q2Falsifying Documents25.4 Sri Lanka2024Q3Falsifying Documents17.3 United Kingdom2023Q4Keeping Undersized fish14.2 Ireland2022Q2Falsifying Documents10.4 United Kingdom2022Q2Prohibited Fishing Gear10.3 a maximum entropy approach. We described each event in terms of an (enforcement country, Illegal fishing sub behavior) tuple. We learn a maximum entropy distribution using iterative proportional fitting (IPF). For each quarter we use incidents from the past year (i.e., past four quarters) to populate the seed matrix which IFP uses to estimate the current quarter. For each cell in the country/sub behavior matrix we then calculated the difference between the estimated and observed, and render it as a z-score. Using|í§| âĽ5 to define a surprising event, we discover 46 surprising country- behavior events, the top 6 seen in fig 2. Examples of these identified surprising incidents are the USA Q2 2022, and LKA 2024 Q3, the first of which involved two Florida men fishing prior to the season starting and catching undersized fish [46], and the second was a case of falsifying tracking data in Saya de Malha, northeast of Madagascar [17]. 4.4 Evaluation For our evaluation set, we had three validators manually label a random selection of 103 source documents, 52 of which were initially classified by IUU+DB as incidents, 25 were related to IUU+ but not involving an incident, and 26 unrelated to IUU+. The validators read through the document and determined the source scope. If the validator found the document to be an incident, they proceeded to classify the incident from the 7 IUU+ types and 41 sub behaviors. Finally the validators extracted the 100 KDEs and enter them into IUU+DB. We log any differences and update the database with the human labeled data. We first evaluate classification performance on the manually validated test set. We report precision, recall, and F1 score for doc- ument scope, IUU+ type, and IUU+ sub-behavior. We use the same metrics to evaluate extraction performance at the field level, and then aggregate the results by KDE group. We also report a mismatch Table 3: Classification Metrics. ClassificationPrecisionRecallF1 Source Scope (macro)0.700.690.64 Incident0.501.000.67 Related, no incident0.640.550.59 Unrelated0.960.520.68 IUU Type (micro)0.820.870.84 IUU Type (macro)0.840.870.82 IUU Sub-Behavior (micro)0.660.810.73 IUU Sub-Behavior (macro)0.740.810.75 Table 4: Extraction Metrics. KDE GroupPrecisionRecallF1Mismatch Rate Species0.710.710.710.86 Products0.670.560.610.88 Compliance0.940.640.760.07 Vessel0.860.480.620.17 Aggregation0.900.360.510.20 Landing1.000.360.530.00 Distribution1.000.360.530.00 Crew0.710.200.310.40 Aquaculture0.830.200.320.33 Event0.390.360.380.64 Labor Standards0.670.320.430.67 Catch0.550.240.330.71 Transshipment0.830.200.320.50 Trade0.670.240.350.60 rate. This measures how often the systemâs extracted value differs from the validated value when both contain a populated field. The classification results in Table 3 show that IUU+DB performs reasonably well across the main classification tasks. However, the confusion matrix, seen in fig 6 and the document-scope results show some weaknesses. For example, incident classification has a precision of 0.50 and a recall of 1.00. This suggests that the classifier is tuned to avoid missing incidents, but it also marks too many doc- uments as incidents. This low precision adds noise to the database and makes downstream analysis more difficult. For sources classified as incidents, we evaluated extraction per- formance; the results are shown in Table 4. We observe four main patterns: high recall with high mismatch, moderate-to-high re- call with low mismatch, low recall with low mismatch, and low- to-moderate recall and high mismatch. In the high-recall, high- mismatch groups, such as species and products, the system often finds the relevant information but also includes noisy or incorrect values. In the moderate-to-high-recall, low-mismatch groups, such as compliance and vessel information, the system performs well and extracts accurate information at a reasonable rate. In the low- recall, low-mismatch groups, such as aggregation, distribution, and landing, the system often misses relevant fields. However, when it does extract them, the values are usually correct. Finally the last group are the fields that the system does not often extract, and when it does it often struggles to do so with high degrees of accu- racy, such as information regarding catch. or labor standards. These IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information ExtractionConferenceâ17, July 2017, Washington, DC, USA Figure 6: Scope Confusion Matrix. Table 5: Performance comparison between IUU+DB, GPT 4o Mini, and GPT 5.4. SourceIUU TypeIUU Sub-Behavior ModelF1macro F1micro F1macro F1micro F1 GPT-4o Mini0.620.700.600.580.53 GPT-5.4 Mini0.620.760.740.620.71 IUU+DB0.650.840.820.730.75 KDE groups highlight areas for improvement with more training examples for prompt optimization. From these statistics we can see that while IUU+DB generally performs well, there is more work needed to reduce the amount of noise, by improving the precision of the scope classifier and reducing the mismatch rate. 4.5 Ablation Comparison against Baselines To evaluate whether IUU+DB improves performance we compare the results of the test set above against two base lines: GPT-4o Mini [31], the model used in our implementation, as well as GPT 5.4 Mini [32] to test how IUU+DB performs w.r.t. newer models. As can be seen from Table 5, all three models perform well. IUU+DB offers a slight improvement on the initial task of classifying the scope against both stand alone LLM calls, but shows 15-20% im- provement over the plain GPT-4o calls and 10% improvement over the newer GPT-5.4 model for IUU+ Type and Sub-Behavior respec- tively. Additionally, we noticed that both GPT-4o Mini and GPT-5.4 Mini extracted about 5 fields fewer than the human extraction or than IUU+DB did; see Fig. 7. 5 Ethical Considerations Intellectual Property. This project aggregates IUU+ content from a range of sources, including some that are behind paywalls or Figure 7: KDEs extracted by source. have restrictions on rehosting under their Terms of Service. When paywalled articles were used, we accessed them through legitimate paid subscriptions or purchased access; the system does not scrape or redistribute protected content without authorization. To respect publisher rights and revenue models, IUU+DB does not show full article text to end users. Instead, users see a summary, source meta- data, and a link to the original document. Full text is retained only on the backend for indexing, validation, and audit purposes, and access is limited to a small set of administrative users. Biased Sources. The source documents that are aggregated reflect the perspectives of the organizations and states that produce them. There may be cases where the illegality of an action is disputed between states, e.g., fishing in disputed waters. This project does not aim to adjudicate these disputes nor claim that all cases within are facts, rather we aim to record claimed instances of IUU+ activity for analysis. Territorial Ambiguity. Some incidents in this dataset occur within territories with contested recognition under international law, such as Somaliland. Assigning an incident to a country necessarily re- quires choices that may favor one political stance over another. For this project, we match locations to ISO-3166 country codes in order to standardize country names. For events that take place within areas that do not have a ISO code we register any location data available and leave the country field empty. 6 Discussion 6.1 Limitations Source Equivalence and Factual Reliability. As previously mentioned, a wide variety of sources are ingested, with a range of credibility. We currently treat all sources as equivalently trustworthy, despite large differences in the evidentiary standards (e.g., a small news organization versus a peer reviewed paper). Users should interpret the incidents as claims of incidents rather than ground truthed fact. Additionally, our data are mostly from English language sources, and refer broadly to events in countries where news is reported in English. Conferenceâ17, July 2017, Washington, DC, USABodwell et al. Data Extraction Quality. IUU+DB is limited by the quality of the source material when determining KDEs like the location and species involved. For example, some sources do not specify any species details at all [8]. Others use generic species group names rather than an exact species distinguisher (e.g., names like âtunaâ, âsalmonâ, âcrabâ, instead of the specific species names such as âyel- lowfin tunaâ, âsockeye salmonâ, âblue swimming crabâ). Meanwhile, other sources include the full scientific name (e.g., âThunnus al- bacaresâ for yellowfin tuna). Incident Bias. The presence of incident reporting should not be the only evaluation criterion used to determine a holistic perspective on the occurrence of IUU+ behaviors. The stories that are covered in the news and through online government agency press releases mean that IUU+ behaviors are being detected and enforcement is taking place. However, in geographies and jurisdictions where en- forcement is lacking or non-existent, there will be few incidents to catalog. Nevertheless, there still may be IUU+ behaviors occurring but the lack of enforcement means they are poorly cataloged in our database. 6.2 Future Work Future work will focus on improving extraction accuracy and strengthening source-scope and IUU+ type classification. Better classification will reduce noise in the database and make it easier for subject matter experts to identify useful patterns and insights. We will also use the generated labels, along with the manually labeled fine-tuning set, to develop and evaluate alternative classification modules that can improve IUU+DBâs performance. We also plan to expand the system to ingest image and figure data. Some important information may appear only in graphics, tables, maps, or other visual formats, and adding this capability would allow IUU+DB to capture a broader range of evidence from each source. Another area of future work is the development of a more user- friendly record validation pipeline that can be opened up to external users. As the database grows, expert validation will remain impor- tant for improving record quality, correcting errors, and helping the system learn from reviewed examples. Finally, we plan to move IUU+DB toward a regularly updated database that can be refreshed on a weekly basis. This will include improved extraction pipelines for each source type and broader coverage of web-scraped sources, including additional NGO, gov- ernment, industry, and media sources from around the world. 7 Conclusion This paper offers IUU+DB as a flexible framework for extracting complex schema from variable length sources at scale. IUU+DB is a tool unique in its scope and potential for near-real time of coverage of IUU+ incidents, giving it the potential to be highly useful to reg- ulators or stakeholder groups. Users can use IUU+DB to quickly compile available recent information on a countryâs or fisheryâs IUU+ incidents, apply risk-based screening to seafood trade con- trols, or better understand links between KDEs (e.g., certain fishing practices and labor abuses) and underlying IUU+ incidents. These varied uses for IUU+DB can help improve fishery conservation and social outcomes. Our results also show that IUU+DB can surface patterns that are difficult to see from individual reports alone. These include the distribution of incidents across IUU+ behavior types, countries, species, and source categories, as well as co-occurrence patterns among IUU+ types and behaviors. Future analyses can use these structured records to identify emerging hotspots, compare reporting patterns across source types, and better understand how different forms of IUU+ activity are connected. This makes IUU+DB not only a database, but also a tool for generating evidence that can inform policy, enforcement priorities, and future research. A Appendix: IUU+ Behaviors and KDE Data Dictionary We define 7 IUU+ types each with a set of related behaviors. They are as follows: 1) Illegal Fishing, constituent behaviors include: exceeding catch quotas, keeping undersized fish, catching unau- thorized, protected or prohibited species, the use of banned or prohibited fishing gear, fishing in closed areas or during closed seasons, obscuring vessel identity, falsifying documents, engag- ing in unauthorized transshipment, engaging in illegal bycatch; 2) Unreported Fishing, constituent behaviors include: unreported or underreported target catch weight or size, unreported or underre- ported discards, bycatch size or weight, misreported target catch species, misreported non target species, misreported gear, misre- ported fishing time or location; 3) Unregulated Fishing, constituent behaviors include: dishing without a flag state, fishing under flag state not party to RFMO, fishing in unregulated areas or for unregu- lated stock; 4) Seafood Fraud or Mislabeling, constituent behaviors include: species mislabeling or fraud, production information fraud; 5) Forced Labor or Labor Abuse, constituent behaviors include: wage and pay violations, abusive living conditions, abusing work- ing conditions, inadequate crew size, physical or sexual violence, workplace intimidation, workers families threatened, isolation, mi- grants threatened; 6) Circumventing Prohibitions or Sanctions, con- stituent behaviors include: circumventing sanctions (individuals or corporations), circumventing import prohibitions (countries or products); 7) Aquaculture Specific Activities, constituent behav- iors include: unapproved or non-native species, illegal sourcing of broodstock or seed, misrepresentation or unauthorized farm opera- tions, unlicensed, unregistered, or unauthorized farm operations, and stolen aquaculture products. We identified the following 14 groups of KDEs: Species, Prod- uct, Event, Vessel, Crew, Labor, Catch, Compliance, Aquaculture, Transshipment, Aggregation, Trade, Distribution, Landing. Taken together these groups have 100 KDE fields. References [1]AFP. 2022.Fish tradeâs murky waters cloud double murder in Ama- zon. https://w.digitaljournal.com/world/fish-trades-murky-waters-cloud- double-murder-in-amazon/article Accessed: 13 May 2026. [2] David J. Agnew, John Pearce, Ganapathiraju Pramod, Tom Peatman, Reg Watson, John R. Beddington, and Tony J. Pitcher. 2009. Estimating the Worldwide Extent of Illegal Fishing. PLOS ONE 4, 2 (Feb. 2009), e4570. doi:10.1371/journal.pone. 0004570 [3] Kathleen Auld, Raphael Baumler, Deukhoon Peter Han, and Francis Neat. 2023. The collective effort of the United Nations Specialised Agencies to tackle the global problem of illegal, unreported and unregulated (IUU) fishing. Ocean & Coastal Management 243 (Sept. 2023), 106720. doi:10.1016/j.ocecoaman.2023. 106720 [4]Dyhia Belhabib. 2026. Spyglass: Global Fishing Crimes Map. http://spyglass.fish/ IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information ExtractionConferenceâ17, July 2017, Washington, DC, USA [5] Benjie B. Talisic, Tito P. Tan. 2022. 3 nabbed for illegal fishing in Bohol. https:// w.sunstar.com.ph/cebu/local-news/3-nabbed-for-illegal-fishing-in-bohol Ac- cessed: 13 May 2026. [6] F. Blaha, A. Vincent, and Y. Piedrahita. 2023. Guidance document: Advancing end-to-end traceability. Critical tracking events and key data elements along capture fisheries and aquaculture value chains. FAO, Rome, Italy. doi:10.4060/c5484en [7] Anthea Dathe, Kiran Hoffmann, and Aline Mangold. 2026. Useful for Ex- ploration, Risky for Precision: Evaluating AI Tools in Academic Research. arXiv:2605.10125 [cs.AI] https://arxiv.org/abs/2605.10125 [8]Environment Agency of the United Kingdom. 2024. 4 licence dodgers receive fines of ÂŁ710 for fishing illegally. https://w.gov.uk/government/news/4-licence- dodgers-receive-fines-of-690-for-fishing-illegally Accessed: 13 May 2026. [9] EU IUU Coalition. 2025. Import control schemes in major seafood markets: a com- parative study of key data elements in the European Union, the United States, Japan and the Republic of Korea. Technical Report. EU IUU Coalition. 33 pages. https: //w.iuuwatch.eu/wp-content/uploads/2025/09/CDS-KDE-Study-FINAL.pdf [10]European Commission. 2026. EU rules to combat IUU fishing. https://oceans- and-fisheries.ec.europa.eu/fisheries/rules/illegal-fishing_en [11]Firstpost.com. 2023. Sri Lanka arrests 14 Indian fishermen for âpoachingâ in its waters, in all 240 so far this year.https://w.firstpost.com/world/sri- lanka-arrests-14-indian-fishermen-for-poaching-in-its-waters-in-all-240-so- far-this-year-13516332.html Accessed: 13 May 2026. [12]Fisheries and Oceans Canada. 2023.Owners of Canadian fishing vessel Ocean Provider fined and over 30,000 pounds of tuna seized. https://w.canada.ca/en/fisheries-oceans/news/2023/09/owners-of- canadian-fishing-vessel-ocean-provider-fined-and-over-30000-pounds- of-tuna-seized.html Accessed: 13 May 2026. [13] GDST. 2026. Global Dialogue on Seafood Traceability. https://thegdst.org/ [14] Sarah M. Glaser, Paige M. Roberts, and Kaija J. Hurlburt. 2019. Foreign Illegal, Unreported, and Unregulated Fishing in Somali Waters Perpetuates Conflict. Frontiers in Marine Science 6 (Dec. 2019). doi:10.3389/fmars.2019.00704 [15] Ryan Y Hodgson, Steven A Robinson, AmĂŠlie C Boutin, Felix K Chan, Joseph R Bennett, Rachel T Buxton, J. Harry Caufield, Dalal E. L Hanna, and Tim Alamen- ciak. 2026. Assessing the effectiveness of ontology-grounded AI term extraction using OntoGPT for environmental evidence synthesis. Environmental Evidence (2026). [16] Gilles E. Hosch and Shelley C. Clarke. 2025. Advancing the Crucial Notion of âInteroperabilityâ in Catch Documentation Schemes. Fisheries Management and Ecology (2025). doi:10.1111/fme.12812 [17]Ian Urbina. 2024. Two men charged with getting early illegal jump on floridas spiny lobster season. https://w.theglobeandmail.com/world/article-saya-de- malha-bank-seagrass-carbon-sink-fishing-threats/ Accessed: 13 May 2026. [18]Island.lk. 2023. Navy detain two Indian trawlers 25 fisherman poaching in Sri Lankan waters.https://island.lk/navy-detain-two-indian-trawlers-25- fishermen-poaching-in-sri-lanka-waters/ Accessed: 13 May 2026. [19] Keith Doucette. 2023. Officers seize $500,000 worth of baby eels outside Halifax amid fishery closure. https://w.thecanadianpressnews.ca/atlantic/officers- seize-500-000-worth-of-baby-eels-outside-halifax-amid-fishery-closure/ article_0d2c90ef-b7b5-5074-bdf6-c2ae340bcbe0.html Accessed: 13 May 2026. [20]Petr Knoth, Drahomira Herrmannova, Matteo Cancellieri, Lucas Anastasiou, Nancy Pontika, Samuel Pearce, Bikash Gyawali, and David Pride. 2023. CORE: A global aggregation service for open access papers. Nature Scientific Data 10, 1 (June 2023), 366. [21] Sha Li, Ayush Sadekar, Nathan Self, Yiqi Su, Lars Andersland, Mira Chaplin, Annabel Zhang, Hyoju Yang, James B Henderson, Krista Wigginton, et al.2025. Exploring LLMs for Scientific Information Extraction Using The SciEx Framework. arXiv preprint arXiv:2512.10004 (2025). doi:10.1186/s13750-026-00381-0 [22]Gloria M. Luque and C. Josh Donlan. 2019. The characterization of seafood mislabeling: A global meta-analysis. Biological Conservation 236 (Aug. 2019), 556â570. doi:10.1016/j.biocon.2019.04.006 [23]Benjamin M. Marshall, Colin T. Strine, Meredith L. Gore, Evan A. Eskew, Oliver C. Stringham, Pedro Cardoso, Sebastian Chekunov, Freyja Watters, Caro- line Fukushima, Pablo GarcĂa-DĂaz, James S. Sinclair, Michael F. Tlusty, Ryan J. Almeida, Jose W. Valdez, and Alice C. Hughes. 2025. Mapping the global dimen- sions of US wildlife imports. Current Biology 35, 16 (Aug. 2025), 3959â3972.e4. doi:10.1016/j.cub.2025.07.012 [24]Channing May-Mavrellis. 2017. Transnational Crime and the Developing World. Technical Report. Global Financial Integrity. 166 pages. https://gfintegrity.org/ report/transnational-crime-and-the-developing-world/ [25] MongaBay. 2026. https://mongabay.org/ Accessed: 2 November 2025. [26] NewsAPI.ai. 2026. Real-time and Historical News Data.https://newsapi.ai Accessed: 2 November 2025. [27] NOAA Fisheries. 2024. Action Plan to Improve the U.S. Seafood Import Monitoring Program. Technical Report. NOAA. 4 pages. https://w.fisheries.noaa.gov/s3/ 2024-11/SIMP-Action-Plan_final.pdf [28]NOAA Fisheries. 2025.Seafood Import Monitoring Program (SIMP). https://w.fisheries.noaa.gov/international/international-affairs/seafood- import-monitoring-program Archive Location: International. [29]NOAA Fisheries. 2026. United States National Oceanic and Atmostsperic Admin- istration National Marine Fisheries Service. https://fisheries.noaa.gov Accessed: 2 November 2025. [30] Oceana. 2026. https://oceana.org Accessed: 2 November 2025. [31] OpenAI, Aaron Hurst, Adam Lerer, Adam P. Goucher, Adam Perelman, and Aditya Ramesh; et al. 2024. GPT-4o System Card. arXiv:2410.21276 [cs.CL] https://arxiv.org/abs/2410.21276 [32]OpenAI, Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, Aidan McLaughlin, and Aiden Low; et al. 2026. OpenAI GPT-5 System Card. arXiv:2601.03267 [cs.CL] https://arxiv.org/abs/2601.03267 [33]Krista Opsahl-Ong, Michael J Ryan, Josh Purtell, David Broman, Christo- pher Potts, Matei Zaharia, and Omar Khattab. 2024.Optimizing In- structions and Demonstrations for Multi-Stage Language Model Programs. arXiv:2406.11695 [cs.CL] https://arxiv.org/abs/2406.11695 [34]Patrick Gouldsbrough. 2022. County Durham man fined for killing fish with spear in River Wear. https://w.thenorthernecho.co.uk/news/20069085.county- durham-man-fined-killing-fish-spear-river-wear/ Accessed: 13 May 2026. [35] Daniel Salas, Abdirahim Sheik Heile, Matthew A. Schnurr, Myles de Jong, Maya Duff, and Kate Swanson. 2026. Seas of unfreedom: A scoping review of labour exploitation as a structural feature of global fisheries. Marine Policy 191 (Sept. 2026), 107147. doi:10.1016/j.marpol.2026.107147 [36] SerpAPI.com. 2025. Google Scholar API.https://serpapi.com/ Accessed: 2 November 2025. [37]Amanda Shaver and Sally Yozell. 2018. Casting a Wider Net: The Security Implications of Illegal, Unreported, and Unregulated Fishing. Technical Report. The Stimson Center. 40 pages.https://ethz.ch/content/dam/ethz/special- interest/gess/cis/center-for-securities-studies/resources/docs/Stimson- %20Casting%20a%20Wider%20Net.pdf [38]John C. Simeone. 2026. U.S. Imports of Fish and Seafood: An Evaluation of Coverage under the Seafood Import Monitoring Program. Technical Report. U.S. IUU Fishing & Labor Rights Coalition. 64 pages. https://w.iuufishing-laborrights.org/ simp-coverage-report [39]S. Spillias, K. M. Ollerhead, M. Andreotta, R. Annand-Jones, F. Boschetti, J. Duggan, D. B. Karcher, C. Paris, R. J. Shellock, and R. Trebilco. 2025. Evaluating generative AI for qualitative data extraction in community-based fisheries management literature. Environmental Evidence (2025). [40] Oliver C. Stringham, Stephanie Moncayo, Eilish Thomas, Sarah Heinrich, Adam Toomes, Jacob Maher, Katherine G. W. Hill, Lewis Mitchell, Joshua V. Ross, Chris R. Shepherd, and Phillip Cassey. 2021. Dataset of seized wildlife and their intended uses. Data in Brief 39 (Dec. 2021), 107531. doi:10.1016/j.dib.2021.107531 [41] Andrew J. Temple, Daniel J. Skerritt, Philippa E. C. Howarth, John Pearce, and Stephen C. Mangi. 2022. Illegal, unregulated and unreported fishing impacts: A systematic review of evidence and proposed future agenda. Marine Policy 139 (2022), 105033. doi:10.1016/j.marpol.2022.105033 [42] TRAFFIC. 2026. Wildlife Trade Portal. https://w.wildlifetradeportal.org/ [43] UN FAO. 2023. Quantifying IUU fishing. UN FAO. doi:10.4060/c6434en [44] UN FAO. 2026. Agreement on Port State Measures (PSMA). https://w.fao. org/port-state-measures/background/en/ [45] UN FAO. 2026. Monitoring, control and surveillance systems to combat illegal, unreported and unregulated fishing. https://elearning.fao.org/course/view.php? id=1128 [46] Undercurrent News. 2022. Two men charged with getting early illegal jump on floridas spiny lobster season. https://w.undercurrentnews.com/2022/06/15/ two-men-charged-with-getting-early-illegal-jump-on-floridas-spiny-lobster- season/ Accessed: 13 May 2026. [47] Undercurrent News. 2024. Mexican drug cartel linked to illegal red snapper harvests in US waters. https://w.undercurrentnews.com/2024/11/28/mexican- drug-cartel-turns-to-lucrative-side-hustle-illegally-harvests-red-snapper-in- us/ Accessed: 13 May 2026. [48]Undercurrent News. 2025. Canadian man gets CAD 1m fine, six years in jail for illegal fishing. https://w.undercurrentnews.com/2025/07/29/canadian-man- gets-cad-1-0m-fine-six-years-in-jail-for-illegal-fishing/ Accessed: 13 May 2026. [49]Undercurrent News. 2026. https://undercurrentnews.com Accessed: 2 November 2025. [50]Undercurrent News. 2026. North Star faces $53,000 fine over Alaska viola- tions. https://w.undercurrentnews.com/2026/02/24/north-star-faces-53000- fine-over-alaska-violations/ Accessed: 13 May 2026. [51]US DOJ. 2026. United States Department of Justice. https://doj.gov Accessed: 2 November 2025. [52] Bruce J. Weissgold. 2024. US wildlife trade data lack quality control nec- essary for accurate scientific interpretation and policy application. Con- servation Letters 17, 2 (2024), e13005.doi:10.1111/conl.13005_eprint: https://conbio.onlinelibrary.wiley.com/doi/pdf/10.1111/conl.13005. [53]WWF. 2014. Illegal Russian Crab: An Investigation of Trade Flow. Technical Report. World Wildlife Fund. 40 pages. https://w.worldwildlife.org/publications/ illegal-russian-crab-an-investigation-of-trade-flow/ [54]Zikang Zhang, Wangjie You, Tianci Wu, Xinrui Wang, Juntao Li, and Min Zhang. 2025. A Survey of Generative Information Extraction. In Proceedings of the Conferenceâ17, July 2017, Washington, DC, USABodwell et al. 31st International Conference on Computational Linguistics, Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, and Steven Schockaert (Eds.). Association for Computational Linguistics, Abu Dhabi, UAE, 4840â4870. https://aclanthology.org/2025.coling-main.324/ Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009