# NuExtract by NuMind > NuExtract is a specialized vision-language model for high-accuracy document extraction: structured JSON extraction and full-document Markdown (OCR/content) extraction. Multimodal, multilingual, deployable privately. NuExtract is built by NuMind. It is a purpose-built LLM that outperforms general-purpose frontier models (GPT-4, Claude, Gemini) for document extraction tasks while running at a fraction of the cost and latency. Available as a hosted API or as a private deployment inside enterprise infrastructure. ## Pages - [Home](/): Platform overview with live extraction demos. - [About](/about): About NuMind — team, mission, and approach. - [Pricing](/pricing): SaaS API usage pricing and enterprise deployment tiers. - [Blog](/blog): Product updates, release notes, and engineering deep-dives. ## Blog posts - [NuExtract3: The Reasoning Open-Source OCR & Structured Extraction LLM](/blog/nuextract-3-release) - [NuExtract Platform: The New Information Extraction](/blog/nuextract-platform-the-new-information-extraction) - [NuExtract 2.0: Outclassing Frontier LLMs in Information Extraction](/blog/outclassing-frontier-llms-nuextract-2-0) - [NuExtract 1.5 — Multilingual, Infinite context, still small, and better than GPT-4o!](/blog/nuextract-1-5-multilingual-infinite-context) - [NuExtract: A Foundation Model for Structured Extraction](/blog/nuextract-a-foundation-model-for-structured-extraction) - [A Foundation Model for Entity Recognition](/blog/a-foundation-model-for-entity-recognition) - [Creating Task-Specific Foundation Models with GPT-4](/blog/creating-task-specific-foundation-models-with-gpt-4) - [What are Large Language Models?](/blog/what-are-large-language-models) - [Seed Round Completed](/blog/seed-round-completed) ## Benchmarks Structured extraction (document to JSON), average score across the NuExtract 3 evaluation suite. Figures transcribed from /blog/nuextract-3/structured-extraction-benchmark.svg. | Model | Parameter size | Non-thinking score | Thinking score | Thinking tokens | | --- | --- | --- | --- | --- | | NuExtract3 | 4B | | 65.2 | 2036 | | Gemma 4 | 4B | | 53.8 | 3005 | | Qwen3.5 | 4B | 50.7 | | 27177 | | GLM 4.6V | Flash (9B) | 48.4 | | 2989 | | Granite-Vision | 4.1-4B | 35.5 | | | | Ministral 3 | 3B (3.8B) | 31.5 | | 27586 | ## Optional - [SaaS Terms & Conditions](/terms/saas) # NuExtract Platform — Structured & OCR Extraction by NuMind https://about.nuextract.ai/ ## World's Best AI for Document Extraction Extract high-accuracy JSON and Markdown with NuExtract, our specialized LLM. Multimodal. Multilingual. Deployable privately. Try NuExtract3 Book a Demo 3 NuExtract Specialized VLM {} JSON Structured output T Markdown Full content output Databases Analytics / BI ERP · CRM RAG Systems Chatbots AI Agents ### Deployed inside the most demanding companies 30% ⟶ 97% Extraction accuracy after switching to NuExtract AXA’s previous automated workflow achieved 30% field-level accuracy, requiring extensive human review. With NuExtract, accuracy increased to 97% (a 23-fold reduction in errors), dramatically lowering manual validation and extraction costs. # About NuMind — The Team Behind NuExtract https://about.nuextract.ai/about ## About NuMind NuMind builds state-of-the-art specialized AI models for structured data extraction & OCR. Our research team brings deep expertise in machine learning and natural language processing, and we published high-impact papers advancing the field. We combine rigorous research with fast product innovation to push the boundaries of what AI can do with respect to document extraction. #### Etienne Bernard Co-Founder & CEO #### Samuel Bernard Co-Founder & CTO #### Francois Laroche Senior Software Engineer #### Alexandre Constantin Machine Learning Scientist #### Olga Chuchuk Software/ML Engineer #### Sören Dréano Machine Learning Engineer #### Mitul Kamal Sales Executive #### Nathan Fradet Machine Learning Scientist #### Shivam Randad Marketing Associate #### Shikha Pandey Marketing Associate Backed by ### Investors & Partners # NuExtract Pricing — SaaS API or Enterprise Deployment https://about.nuextract.ai/pricing ## NuExtract Pricing — SaaS or Enterprise Usage-based access for the API, tiered pricing for private enterprise deployment. ### NuExtract API Usage-based API access. Input tokens $1 /M Output tokens $5 /M Onboarding Instant Volume discounts Available Start for Free ### NuExtract Enterprise Private deployment for enterprise environments. Deployment Private cloud / on-prem Confidentiality By-design Fine-tuning Optional Pricing Tiered usage Talk to Sales ### NuExtract API: Usage-based pricing We charge usage of the model when extracting information from documents. The price depends on the number of input tokens (template, examples, and input document) and output tokens (JSON output). Model NuExtract 3 Input tokens $1 /M Output tokens $5 /M Token (Image) 32×32 pixels Token (Text) word or sub-word Model Input tokens Output tokens Token (Image) Token (Text) NuExtract 3 $1 /M tokens $5 /M tokens 32×32 pixels word or sub-word Processing large volumes? Per-token prices can be significantly lowered through batching or fine-tuned models. Talk to us . #### Estimating token numbers Text tokens In English, 1 word ≈ 1.3 tokens on average — a page typically contains ~1,000 tokens . Some languages have higher token counts per word. Image tokens With NuExtract 3, a token corresponds to a 32×32 pixel patch . An A4 page at 115 dpi ≈ 1,500 tokens. Text tokens In English, 1 word ≈ 1.3 tokens on average — a page typically contains ~1,000 tokens . Some languages have higher token counts per word. Estimated cost for 1 and 100 text pages. View estimates 1 page ~$0.001 ~1,000 input tokens 100 pages ~$0.10 ~100k input tokens Image tokens With NuExtract 3, a token corresponds to a 32×32 pixel patch . An A4 page at 115 dpi ≈ 1,500 tokens. Estimated cost for 1 and 100 scanned A4 pages. View estimates 1 A4 page @ 115dpi ~$0.0015 ~1,500 input tokens 100 A4 pages @ 115dpi ~$0.15 ~150k input tokens Modality Size Input tokens Input price Text 1 page ~1,000 ~$0.001 Text 100 pages ~100k ~$0.10 Image 1 A4 page @ 115dpi ~1,500 ~$0.0015 Image 100 A4 pages @ 115dpi ~150k ~$0.15 PDFs and other formatted documents are converted to images by default. ### Frequently asked questions #### Can I run NuExtract on my own infrastructure? #### Do you offer volume discounts on the API? #### Is fine-tuning available? #### How are tokens counted for documents? # SaaS Terms & Conditions — NuExtract API Services https://about.nuextract.ai/terms/saas Legal ## Terms and Conditions — NuExtract API Services Date of entry into effect: July 20, 2026 ### Recitals These Terms and Conditions (the “Terms” ) govern the access to and use of the NuExtract API services (the “Services” ), provided by NuMind Technology, Inc., a Delaware corporation with offices at 2093 Philadelphia Pike, #3795, Claymont, DE 19703, United States of America ( “NuMind” ). These Terms apply to each (i) online registration completed by the Customer through the Website, or (ii) quotation, order form, purchase order, or similar commercial document (each, an “Order Form” ), agreed or accepted by the Customer for the access to and use of the Services. By completing an online registration for the Services on the Website, signing or accepting an Order Form, accessing the API, or by issuing a purchase order referencing these Terms, the Customer agrees to be bound by: (i) these Terms; (ii) the applicable Order Form (where applicable); (iii) the Data Processing Agreement ( “DPA” ) as further described in Section 12; and (iv) any additional specific conditions or statements of work agreed between the Parties. Together, these documents form the “Agreement” . The Services are intended for professional use only. ### 1. Definitions For the purposes of these Terms, the following words, whether used in the singular or plural, shall have the meanings set out below: “Affiliate” means, if and when applicable, any present or future legal entity controlled by the Customer, it being specified that Control shall mean owning directly or indirectly more than 50% of the voting rights or of the shares or interests of said Affiliate, or deemed under the Customer's control under the applicable law. The provisions herein are applicable to the Affiliates, and the Customer guarantees and ensures compliance to these Terms by its Affiliates. “Agreement” has the meaning set forth in the Introduction. “API” means NuMind's NuExtract application programming interface through which the Customer and its Authorized Users access the Services. “API Credentials” means the API keys, access tokens, or other authentication credentials issued by NuMind to the Customer and its Authorized Users for the purpose of accessing the API. “Authorized User” means any natural person duly authorized by the Customer to access and use the Services in accordance with these provisions, on the Customer's behalf and under the Customer's sole responsibility. The Customer guarantees and ensures compliance with the Agreement by all its Authorized Users. “Business Day” means any day other than a Saturday, Sunday, or a public holiday in the jurisdiction of NuMind. “Confidential Information” means any information disclosed by one party to the other party, in any form, that is marked as confidential or that reasonably should be understood to be confidential given its nature or the circumstances of disclosure as defined in Section 11. “Customer Content” means all data, documents, files, images, text, prompts, schemas, templates, extraction instructions, and other inputs or materials submitted by or on behalf of the Customer to the API or to NuMind in connection with the Services, excluding Outputs. “Documentation” means NuMind's official user guides, API reference documentation, technical specifications, and similar written materials that NuMind makes available for the Services, including materials available on the Website and the Pricing Page, as updated from time to time, excluding marketing materials, sales presentations, roadmaps, and oral statements unless expressly incorporated into the Agreement. “DPA” means the Data Processing Agreement entered into or to be entered into between the parties pursuant to Section 12. “Effective Date” means the earlier of: (i) the date on which the first Order Form is signed or otherwise accepted; or (ii) the date on which the Customer first accesses the Services. “Fees” means the charges payable by the Customer for access to and use of the Services, calculated on the basis of Token consumption as further described in Section 9. “Order Form” means the commercial document or online acceptance mechanism by which the Customer subscribes to the Services, specifying the applicable Services, Fees, usage parameters, Term, and any specific commercial terms, including: (i) an online registration or subscription completed by the Customer through the Website, or (ii) any quotation, order form, or purchase order signed, issued, or accepted by the Customer and referencing these Terms. Where the Customer subscribes online, the online registration confirmation or API Credentials provisioning notification from NuMind shall constitute the Order Form for the purposes of the Agreement. “Output” means the result generated by the Services from Customer Content in response to an API request, including structured data, extracted content, text, markdown, or similar results. “Pricing Page” means NuMind's applicable pricing page for the Services, available as of the date of these Terms at https://about.nuextract.ai/pricing , or any successor URL communicated by NuMind in accordance with Section 17, setting out the applicable reference Token rates, pricing methodology, and available volume discount options. In the event of any conflict between the Pricing Page and an applicable Order Form, the Order Form shall prevail. “Services” means NuMind's NuExtract API services, comprising an AI-powered application programming interface enabling the extraction of structured information from documents, as further described in the applicable Order Form. “Token” means the unit of measurement used to quantify API usage for billing purposes. Token consumption for billing purposes is further described on the Pricing Page or where applicable an Order Form. “Website” means NuMind's official website(s) and online platform(s), including https://numind.ai and https://about.nuextract.ai , through which NuMind makes available information about the Services, the Documentation, the Pricing Page, and the online subscription or registration mechanism for the Services. ### 2. Scope and Acceptance 2.1 These Terms govern the conditions of access to and use of the Services by the Customer and its Authorized Users. The Services are reserved for professionals acting exclusively in the context of their professional activity. 2.2 By signing or otherwise accepting an Order Form, by completing an online registration for the Services on the Website, by accessing the API or the Services, or by issuing a purchase order referencing these Terms, the Customer acknowledges having read and fully accepts these Terms. 2.3 These Terms are made available to the Customer prior to or upon registration to the Services. In the event of any contradiction between these Terms and any general terms and conditions of purchase or any other general or special terms and conditions issued by the Customer, these Terms shall prevail. 2.4 NuMind may modify these Terms at any time. In the event of any material modification, NuMind shall notify the Customer by email or, where applicable, any other means agreed by the Parties, at least thirty (30) days before such modifications take effect. The Customer's continued access to or use of the Services after the expiry of such notice period shall constitute the Customer's full acceptance of the modified Terms. If the Customer does not accept the proposed modifications, it may terminate the Agreement in accordance with Section 16 before the end of such notice period. The conditions applicable to modifications of Terms during an ongoing Order Form, including the Customer's options upon objection, are further described in Section 17. 2.5 NuMind reserves the right to update, modify, improve, or discontinue the Services or any feature or functionality thereof at any time, subject to the notice provisions set out in Section 17.3. NuMind may offer additional services. These may be subject to additional, distinct or supplementary terms and conditions, as well as additional financial terms and quote, if and where appropriate. ### 3. Access to the Services and API Credentials 3.1 Access to the Services is provided via the API. NuMind shall make API Credentials available to the Customer in accordance with the applicable Order Form, for the purpose of authenticating API calls and accessing the Services. 3.2 The Customer may have access to an online administrative interface allowing the Customer to manage its API Credentials and monitor its Token consumption. 3.3 The Customer is solely responsible, with respect to NuMind, for the security and confidentiality of its API Credentials. The Customer shall ensure that API Credentials are not disclosed to unauthorized persons and are used solely in accordance with the Agreement. 3.4 The Customer shall notify NuMind immediately upon becoming aware of any actual or reasonably suspected unauthorized use of its API Credentials or of any security breach relating to the Services. 3.5 Access to the Services via the API requires that the Customer maintains adequate technical infrastructure, systems, and Internet connectivity. NuMind shall not be liable for difficulties in accessing the Services arising from the Customer's own network, equipment failures, or Internet connectivity disruptions. 3.6 NuMind shall not be liable for any unauthorized access to or use of the Services resulting from the Customer's failure to adequately secure its API Credentials in accordance with this Section. ### 4. License Grant 4.1 Subject to the Customer's compliance with the Agreement and payment of all applicable Fees, NuMind grants the Customer, for the duration of the Agreement, a limited, non-exclusive, non-transferable, non-sublicensable right to access and use the Services via the API, solely for the Customer's internal business purposes. 4.2 The Customer may authorize its Authorized Users to access and use the Services on its behalf, solely in accordance with the Agreement and under the Customer's sole responsibility. 4.3 No rights are granted other than those expressly stated in the Agreement. In particular, no right is granted to access source code, to access or export NuMind's underlying AI models separately from the Services. ### 5. Restrictions on Use The Customer shall not, and shall not permit any Authorized User or third party to: 5.1 sell, resell, license, sublicense, rent, lease, distribute, transfer, assign, or otherwise make the Services or the API available to any third party as a standalone product or service; 5.2 reverse engineer, decompile, disassemble, translate, or otherwise attempt to derive the source code, underlying model weights, architecture, or algorithms of the Services or NuMind's AI models, except to the extent expressly required by mandatory applicable law; 5.3 perform or facilitate model extraction, model theft, reconstruction of model weights or components, adversarial prompting intended to extract model internals, or any other attempt to access, replicate, or approximate the architecture or internals of NuMind's AI models; 5.4 remove, alter, or obscure any proprietary notices, technical protection measures, or security controls embedded in the Services or the API; 5.5 use the Services beyond the scope permitted under the Agreement or the applicable Order Form; 5.6 use the Services in violation of applicable law; 5.7 use NuMind's Confidential Information, the Documentation, model Outputs, API responses, or other non-public technical information concerning the Services to develop, train, fine-tune, or commercialize a competing document extraction, information extraction, or AI data extraction product or service; 5.8 circumvent, disable, or tamper with any access controls, rate-limiting mechanisms, Token counting systems, or usage-tracking mechanisms of the Services; 5.9 use the Services to process Customer Content that is unlawful, fraudulent, or infringes third-party rights; 5.10 To avoid all doubt, the Customer may freely use Outputs for any lawful business purpose, including in its own products, services, and internal workflows, subject to terms of Section 7. ### 6. Customer Obligations and Responsibilities 6.1 The Customer is solely responsible for: (a) the acts and omissions of its Authorized Users; (b) the security and management of its API Credentials in accordance with Section 3; (c) ensuring that Customer Content and all use of the Services comply with applicable law and the Agreement; (d) obtaining and maintaining all necessary rights, permissions, consents, and legal bases for Customer Content, including for the processing of any personal data contained therein; (e) reviewing and validating all Outputs before relying on them; and (f) providing NuMind with all information and cooperation reasonably necessary for NuMind to perform the Services. 6.2 The Customer is responsible for determining and complying with all laws and sector-specific requirements applicable to its Customer Content and use of the Services. Nothing in this Section limits NuMind's obligations under the Agreement or applicable law. 6.3 The Customer is responsible for ensuring a sufficient level of AI literacy and technical understanding among its Authorized Users regarding the capabilities and limitations of the Services, taking into account their technical background, training, and the context in which the Services are used. 6.4 The Customer guarantees and ensures compliance with the Agreement by all its Authorized Users and remains fully liable to NuMind for any breach of the Agreement by its Authorized Users. ### 7. Customer Content, Outputs and Data #### 7.1 Ownership The Customer retains all right, title, and interest in and to Customer Content. The Customer retains all right, title, and interest in and to Outputs, subject to NuMind's ownership of the Services, the underlying AI models, and related technology as described in Section 10. #### 7.2 Limited License to NuMind The Customer grants NuMind a limited, non-exclusive, royalty-free, worldwide license to access, host, copy, and process Customer Content solely to the extent strictly necessary and for the sole purpose of providing the Services to the Customer under the Agreement. This license does not extend beyond what is strictly necessary to perform the Services and terminates automatically upon deletion of Customer Content pursuant to Section 7.3. #### 7.3 Data Deletion 7.3.1 NuMind shall automatically delete all Customer Content and Outputs from its active production systems within a maximum of fourteen (14) days following the relevant API call. 7.3.2 Residual copies may remain in encrypted database backup or history systems for a maximum of thirty (30) days. Such copies are not used for ordinary processing and are deleted through the ordinary backup rotation process. 7.3.3 NuMind provides the Customer with a dedicated API endpoint enabling the Customer to request the deletion of Customer Content and Outputs at any time. Following a valid deletion request, NuMind shall delete the relevant Customer Content and Outputs from its active production systems without undue delay. Residual copies may remain until the expiry of the backup retention period described in Section 7.3.2. 7.3.4 NuMind shall not retain, aggregate or store Customer Content or Outputs except as necessary to provide and secure the Services, comply with applicable law, and maintain the limited backup copies described in Section 7.3.2. #### 7.4 No Training NuMind shall not use Customer Content, Outputs, or any data submitted by the Customer through the Services for any of the following purposes: (a) training, fine-tuning, benchmarking or improving NuMind's AI models; (b) developing, improving, or commercializing new or existing products or services; (c) any other purpose beyond the performance of the Services for the Customer under the Agreement. #### 7.5 Customer Responsibility for Customer Content The Customer is solely responsible for the accuracy, completeness, lawfulness, quality, and suitability of Customer Content. NuMind shall not be liable for any inaccuracies, defects, or unlawful content in Customer Content, or for any consequences arising from the use of Customer Content that does not comply with applicable law. #### 7.6 AI Outputs and Limitations 7.6.1 The Customer acknowledges that Outputs generated by the Services may be incomplete, inaccurate, inconsistent, or otherwise unsuitable for certain purposes. NuMind does not warrant that Outputs will be accurate, complete, error-free, or fit for any particular purpose. 7.6.2 The Customer acknowledges and agrees that: (i) artificial intelligence systems, by their nature, may produce errors, omissions, inaccuracies, hallucinations, or unexpected results; (ii) the accuracy, completeness, and reliability of Outputs depend on various factors including the quality of Customer Content, document type and format, extraction schema complexity, and the inherent probabilistic nature of AI models; and (iii) Outputs are provided for support purposes only and do not constitute advice, recommendations, guarantees, or professional opinions of any kind. 7.6.3 The Customer is solely responsible for reviewing and validating all Outputs before relying on them for any purpose, including any decision, action, filing, communication, or any decision producing legal effects concerning a natural person. #### 7.7 Regulatory Use The Customer is responsible for determining whether its specific use of the Services or Outputs is subject to AI-specific, sector-specific, or other legal or regulatory requirements, and for implementing all required human oversight measures, controls, risk management procedures, recordkeeping obligations, and compliance safeguards applicable to its specific use case. ### 8. Service Availability, Maintenance and Support #### 8.1 Availability of the Services The Services are accessible 24 hours a day and 7 days a week, except in the event of an interruption, scheduled or unscheduled, for maintenance purposes or in the event of force majeure. NuMind shall make its best efforts to inform the Customer prior to the carrying out of maintenance operations or updates. NuMind shall not be liable in relation to such operations. NuMind may update and improve the Services from time to time. NuMind shall implement appropriate technical and organisational measures designed to protect the Services, as described in the DPA. Thus, in the event of interruption of service and whatever the cause, NuMind shall make its best efforts in order for the Services to be put back into service as soon as possible. NuMind reserves the right to interrupt the operation of the Services or to prohibit the access to the Services when the security of the Services is threatened (security flaw detected, intrusion, data corruption, virus, malware). NuMind may also carry out planned shutdowns of Services, in part or in whole, in particular to carry out maintenance work or updates of the Services. These shutdowns and maintenance work shall be carried out as far as possible during periods of low activity. NuMind shall make in this case its best efforts, when possible, to notify the Customer in advance of any planned shutdown of Services. NuMind undertakes to restore as soon as possible the access to the Services. NuMind is bound by a best-efforts obligation. NuMind shall not in consequence be liable for any direct or indirect damage suffered by the Customer and/or its Authorized Users resulting from the unavailability of the Services, in whole or in part, and no credit note, refund or credit in any form whatsoever shall be emitted in the event of a shutdown under the terms of this Article. #### 8.2 Updates and Modifications NuMind manages, deploys, and updates the Services as SaaS infrastructure without requiring action by the Customer unless otherwise notified. NuMind reserves the right to update, modify, improve, or discontinue any feature or functionality of the Services at any time. Where a material change to the API or the Services would require the Customer to update its technical integration, NuMind shall use commercially reasonable efforts to provide advance notice in accordance with Section 17. #### 8.3 Support NuMind shall provide reasonable technical support to the Customer for issues relating to the use of the Services. Support is available via the channels designated by NuMind in the applicable Order Form or Documentation. Support is provided on a commercially reasonable efforts basis. #### 8.4 NuMind shall not be liable for any disruption, degradation, dysfunction, or impossibility of access to the Services caused by: (a) factors outside NuMind's reasonable control, including congestion of the Internet network or any other external cause; (b) the Customer's own infrastructure, systems, Internet connectivity, or equipment (including equipment that is not adapted to access the Services); (c) disturbances attributable to the Customer's or any Authorized User's access provider; (d) third-party service providers not under NuMind's control; or (e) a cyber-attack, malicious act, or other security incident directed at the Customer's systems. ### 9. Fees, Invoicing and Payment #### 9.1 Token-Based Pricing 9.1.1 In consideration for access to and use of the Services, the Customer shall pay Fees based on its Token consumption. Fees are calculated based on the aggregate number of Input Tokens and Output Tokens consumed through the Customer's API calls during the relevant billing period, at the rates set forth in the applicable Order Form or, in the absence of a specific Order Form, at NuMind's then-current published rates. 9.1.2 The reference rates for the NuExtract API are defined in the Order Form and/or the Pricing Page. Volume discounts may be available as agreed in the applicable Order Form. Current pricing, indicative cost estimates, and volume discount options are also published on the Pricing Page. 9.1.3 All Fees are defined in United States Dollars (USD) unless otherwise agreed in the applicable Order Form. #### 9.2 Invoicing and Payment 9.2.1 NuMind shall issue invoices on a monthly basis, calculated on the Customer's aggregate Token consumption during the relevant billing month, unless the applicable Order Form specifies a different billing frequency or method. 9.2.2 Invoices are payable within thirty (30) calendar days from the invoice date, by wire transfer (bank transfer), credit card, or via third-party payment processors designated by NuMind from time to time. The available payment methods shall be communicated to the Customer by NuMind in the applicable Order Form or on NuMind's Website. Payment methods and providers are subject to evolution. 9.2.3 The Customer shall pay the total amount of each invoice, all taxes included where applicable, and may not operate any compensation with any sums due or claimed to be due by NuMind. The Customer agrees to pay all taxes, government fees, transfer fees and all other taxes applicable to all payments made. Any bank charges or fees or such from any other intermediaries related to the payment or any incident shall be borne exclusively by the Customer. The Customer undertakes that all sums paid by the Customer shall be of the amount provided for herein without deduction by the Customer of any amount such as any local tax or withholding tax, which shall be solely borne by the Customer. The Customer shall comply with all its obligations of payment and accepts that NuMind retains if required and necessary the aforementioned information of payment according to the conditions and applicable legal durations. The Customer also agrees to the following: - Following the place of transaction, exchange transaction fees or different prices (for example, exchange rates) may be applicable. - Where automatic payment is enabled, the Customer's designated payment method may be charged when an invoice is issued. 9.2.4 In the event of a dispute relating to an invoice, payment of the disputed amount remains due pending resolution. If the dispute is accepted, a credit note shall be issued and provided to the Customer promptly. #### 9.3 Late Payment In the event of late payment of any undisputed amount, late-payment penalties shall accrue automatically, without prior formal notice, from the day following the due date. Late-payment penalties shall accrue at a rate of ten percent (10%) per annum of the overdue amount, calculated from the due date until the date of actual payment in full. If any undisputed amount remains unpaid fifteen (15) calendar days after written notice from NuMind, NuMind may suspend the Customer's access to the Services until full payment, including all accrued late-payment penalties, is received, without prejudice to NuMind's right to terminate the Agreement in accordance with Section 16 and to claim any damages arising from such non-payment. #### 9.4 Price Modifications NuMind reserves the right to modify its Fees and Token pricing rates at any time, subject to prior written notice to the Customer in accordance with Section 17. Modified pricing shall take effect on the date specified in the notice. If the Customer does not accept the modified pricing, it may terminate the Agreement in accordance with Section 16 before the new pricing takes effect, without any penalty for such termination solely on this basis. #### 9.5 No Refunds All Fees paid or due are non-refundable and non-cancellable. In particular, Fees for Tokens already consumed are non-refundable under any circumstances, including in the event of termination. ### 10. Intellectual Property #### 10.1 NuMind's Intellectual Property NuMind and where or if applicable its licensors retain all right, title, and interest in and to the Services, the API, the underlying AI models, model components, algorithms, training pipelines, software, source code, Documentation, methods, generic know-how, updates, improvements, and all related intellectual property rights worldwide. No ownership, title, or other right in NuMind's intellectual property is transferred to the Customer by the Agreement. #### 10.2 Customer's Intellectual Property The Customer retains all right, title, and interest in and to Customer Content and Outputs, subject to the terms of Section 7 and to NuMind's ownership of the Services and underlying technology. The limited license granted to NuMind under Section 7.2 does not constitute a transfer of any intellectual property right of the Customer. #### 10.3 Restrictions The Customer undertakes not to reproduce, adapt, modify, decompile, disassemble, distribute, exploit, or otherwise make available, directly or indirectly, any element of the Services or the API beyond the use expressly authorized under the Agreement. Any unauthorized use of the Services shall be deemed an infringement actionable under applicable intellectual property law. #### 10.4 Feedback Any feedback, suggestions, ideas, or improvement requests provided by the Customer regarding the Services may be used by NuMind freely, without restriction, compensation, attribution, or obligation to the Customer, provided that NuMind shall not acquire any right in Customer Content as a result. #### 10.5 IP Indemnification 10.5.1 If a third party claims that the Services, as provided by NuMind and used by the Customer in accordance with the Agreement, infringe such third party's intellectual property rights, NuMind shall, at its option and as the Customer's sole and exclusive remedy for such claim: (i) defend such claim and pay any damages finally awarded by a court of competent jurisdiction or amounts agreed in settlement by NuMind; and (ii) where commercially reasonable: (A) procure the right for the Customer to continue using the affected portion of the Services; (B) modify or replace the affected portion so that it becomes non-infringing without materially reducing its functionality; or (C) if neither (A) nor (B) is commercially reasonable, terminate the affected portion of the Agreement on written notice; provided that the Customer: (i) promptly provides NuMind with written notice of the claim; (ii) grants NuMind sole control of the defense and settlement; and (iii) provides reasonable cooperation at NuMind's expense. NuMind shall not settle any claim in a manner that admits fault of, or imposes payment or material obligation on, the Customer without the Customer's prior written consent. 10.5.2 Section 10.5.1 does not apply to any claim arising from: (i) Customer Content, Customer instructions, or Customer-provided specifications; (ii) combination of the Services with products, services, or data not provided by NuMind where the claim would not have arisen but for such combination; (iii) modifications to the Services not made or authorized by NuMind; (iv) use of the Services outside the scope of the Agreement or contrary to the Documentation; or (v) the Customer's continued use of the Services after NuMind has offered a non-infringing alternative under Section 10.5.1. ### 11. Confidentiality 11.1 Each party (the “Receiving Party” ) shall keep the other party's (the “Disclosing Party” ) Confidential Information strictly confidential and shall use it only for the purposes of performing or exercising its rights under the Agreement. “Confidential Information” means any non-public information disclosed by one party to the other, in any form, that is marked as confidential or that a reasonable party would understand to be confidential given the nature of the information or the circumstances of disclosure. 11.2 Neither party shall disclose the other party's Confidential Information to any third party except to its employees, individual contractors, advisers, auditors, insurers, or sub-processors who have a legitimate need to know it for the purposes of the Agreement and who are bound by confidentiality obligations no less protective than those set out in this Section. 11.3 The obligations in this Section do not apply to information that the Receiving Party can demonstrate: (a) is or becomes publicly available without breach of the Agreement; (b) was already lawfully known to the Receiving Party without confidentiality obligation at the time of disclosure; (c) is lawfully received from a third party without confidentiality obligation; or (d) is independently developed by the Receiving Party without reference to or use of the Disclosing Party's Confidential Information. 11.4 A party may disclose Confidential Information to the extent required by applicable law, regulation, court order, or competent authority, provided that, where legally permitted, it provides prior written notice to the other party and takes all reasonable steps to minimize the scope and impact of such disclosure. 11.5 The confidentiality obligations in this Section survive expiry or termination of the Agreement for a period of three (3) years, provided that trade secrets shall remain protected for as long as they remain trade secrets under applicable law. ### 12. Personal Data 12.1 Each party shall comply with all applicable data protection and privacy laws in connection with the Agreement. 12.2 In connection with personal data contained in Customer Content, the Customer acts as a controller or processor, as applicable, and NuMind acts as a processor or sub-processor, as applicable. NuMind shall process such personal data only on documented instructions and only to provide the Services, subject to Section 7.3 and the DPA. 12.3 The parties agree upon a Data Processing Agreement ( “DPA” ) governing the conditions under which NuMind processes personal data on behalf of the Customer, and which is deemed incorporated to these Terms. NuMind's standard DPA is available upon request, and in any case on the Website. The Customer may also propose an alternative DPA, subject to NuMind's prior written agreement. 12.4 The Customer is solely responsible for: (a) determining whether and on what legal basis personal data is included in Customer Content; (b) providing all required privacy notices to data subjects; (c) obtaining all required consents and ensuring all applicable legal bases for the processing of personal data contained in Customer Content; and (d) ensuring compliance with all applicable data protection law in connection with its use of the Services. 12.5 The Customer generally authorizes NuMind to use affiliates and subprocessors for the performance of the Agreement. NuMind shall impose on such subprocessors data protection obligations no less protective than those set out in the Agreement and shall remain responsible for their performance. NuMind shall maintain and make available to the Customer, upon request, an up-to-date list of sub-processors that may access or process Customer Content in connection with the Services. NuMind shall notify the Customer in advance of any addition or replacement of sub-processors in accordance with the terms of the DPA. ### 13. Warranties and Disclaimer 13.1 Each party represents and warrants that it has the full legal authority and capacity to enter into and perform its obligations under the Agreement. 13.2 NuMind warrants that, during the term of the Agreement, the Services, as provided by NuMind and when used in accordance with the Agreement and the Documentation, will substantially conform to the Documentation. 13.3 The Customer's exclusive remedy for a material breach of the warranty in Section 13.2 shall be for NuMind to use commercially reasonable efforts to correct or replace the non-conforming portion of the Services within a reasonable time after receiving written notice from the Customer. If NuMind fails to do so within a reasonable time, the Customer may terminate the affected portion of the Agreement. 13.4 Except as expressly stated in the Agreement, the Services are provided on an “as is” and “as available” basis, to the maximum extent permitted by applicable law: (a) NuMind does not warrant that Outputs will be accurate, complete, error-free, or fit for any particular purpose; (b) NuMind does not warrant uninterrupted or error-free availability of the Services; (c) open-source and third-party components incorporated in the Services are provided subject to their respective licenses and, to the maximum extent permitted by law, on an “as is” basis; (d) NuMind does not warrant that the Services are suitable to be used as the sole basis for decisions producing legal effects concerning natural persons or similarly significantly affecting them; (e) to the maximum extent permitted by applicable law, NuMind expressly disclaims any warranty, express or implied, including any implied warranty of merchantability, fitness for a particular purpose, or non-infringement. NuMind does not warrant that the Services will operate error-free or uninterrupted, or that all defects can be or will be corrected. The Customer acknowledges that software products inherently contain errors and that not all errors are economically or technically rectifiable. ### 14. Liability 14.1 NuMind's liability relating to the provision of the Services is a best-efforts obligation. 14.2 To the maximum extent permitted by applicable law, neither party shall be liable to the other for any indirect, incidental, special, consequential, punitive, or exemplary damages, including loss of profits, loss of revenue, loss of business, loss of goodwill, loss of anticipated savings, or loss of data, even if advised of the possibility of such damages. 14.3 To the maximum extent permitted by applicable law, each party's total aggregate liability arising out of or in connection with the Agreement shall not exceed the total Fees paid by the Customer to NuMind during the twelve (12) calendar months preceding the event giving rise to the claim. 14.4 Each party's total aggregate liability for breach of Section 11 (Confidentiality) or Section 12 (Personal Data) shall not exceed two (2) times the cap set out in Section 14.3, subject to the maximum extent permitted by applicable law. Nothing in this Section limits liability to the extent that applying such limitation would contradict Clause 12 of the applicable Standard Contractual Clauses or applicable mandatory law. 14.5 The liability limitations in this Section do not apply to: (a) the Customer's payment obligations under Section 9; (b) the Customer's breach of Section 5 (Restrictions on Use) to the extent consisting of: unauthorized sublicensing or redistribution of the Services; model extraction or theft prohibited by Sections 5.2 and 5.3; use of NuMind's Confidential Information to develop a competing product in breach of Section 5.7; or use of the Services after expiry or termination of the Agreement; and (c) fraud, willful misconduct, or gross negligence, to the extent liability for such cannot be excluded or limited under applicable mandatory law. 14.6 NuMind shall not be liable for any loss or damage directly resulting from: (a) contamination of the Customer's systems by viruses, malicious code, or cyber-attacks or malicious acts of third parties, provided NuMind has implemented the security measures set forth in the DPA; (b) use of the Services in any manner not expressly authorized under the Agreement; (c) continued use of the Services after NuMind has issued a written recommendation to suspend use in connection with a known defect or security vulnerability; (d) use of the Services in an environment or configuration that does not meet NuMind's technical requirements as set out in the Documentation; or (e) use of the Services after expiry or termination of the Agreement. 14.7 The Customer shall procure that all Authorised Users comply with the Agreement and shall remain liable to NuMind for any breach of the Agreement by its Authorised Users. The Customer is responsible for the accuracy, legality, and quality of the Customer Content it provides under the Agreement, and for ensuring that such Customer Content is submitted in compliance with applicable law. Nothing in this Section 14.7 affects NuMind's obligations as a processor of personal data under Section 12 and the applicable DPA. In the event of any breach of the Agreement by the Customer or its Authorised Users, NuMind may exercise any of its rights and remedies under the Agreement, including suspension of access pursuant to Section 16.3 and termination pursuant to Section 16.2 without prejudice to any claim for damages or injunctive relief available under applicable law. 14.8 Notwithstanding the exclusive jurisdiction provisions of Section 18, nothing in the Agreement prevents either party from seeking injunctive or other equitable relief from any court of competent jurisdiction in the event of unauthorized use of its intellectual property or unauthorized disclosure of its Confidential Information, without prejudice to any other rights or remedies under the Agreement or applicable law. 14.9 Subject to mandatory applicable laws, any claim or action arising out of or in connection with the Agreement must be commenced within two (2) years from the date on which the party bringing such claim knew, or reasonably should have known, of the facts giving rise to it. The shortened limitation period set out in this Section 14.9 shall not apply to: (a) claims arising out of fraud or wilful misconduct; (b) claims arising out of a party's obligations under applicable data protection law, including Regulation (EU) 2016/679 (GDPR); or (c) claims arising out of a party's obligations under applicable artificial intelligence regulation, including Regulation (EU) 2024/1689 (AI Act). ### 15. Artificial Intelligence The following provisions apply to the extent that the AI Act is applicable to either party pursuant to Article 2 of the AI Act: #### 15.1 Qualification NuMind is the provider of the Services within the meaning of Regulation (EU) 2024/1689 (the “AI Act” ). The Customer is the deployer of the Services within the meaning of the AI Act when using the Services for its own purposes or those of its clients. #### 15.2 Intended Purpose The intended purpose of the Services is AI-powered structured data extraction from documents, as described in the Documentation. The Customer shall not use the Services for any purpose materially different from the intended purpose without NuMind's prior written consent. Any material change in intended use may alter the risk classification of the Services and the applicable regulatory obligations of each party. #### 15.3 Risk Classification NuMind shall inform the Customer in writing of the risk classification of the Services under the AI Act (including whether the Services constitute a high-risk AI system) as reasonably determined by NuMind. NuMind shall notify the Customer without undue delay if the risk classification changes following a material update to the Services, the Documentation, or applicable law. #### 15.4 NuMind Provider Obligations As provider of the Services, NuMind is responsible for: (i) establishing and maintaining the quality management system required under Article 17 of the AI Act (if applicable); (ii) ensuring the conformity assessment required under Article 43 of the AI Act (if applicable); (iii) registering the Services in the EU AI database under Article 71 of the AI Act (if applicable); (iv) providing the Customer with adequate instructions for use, including information on the capabilities and limitations of the Services, performance metrics, and recommended human oversight measures; (v) establishing and maintaining a post-market monitoring system in accordance with Article 72 of the AI Act (if applicable), taking into account information provided by the Customer pursuant to Section 15.5(viii); (vi) notifying the relevant market surveillance authority of any serious incident in accordance with Article 73 of the AI Act (if applicable); and (vii) retaining the technical documentation for the Services for the period required under Article 18 of the AI Act (if applicable). #### 15.5 Customer Deployer Obligations As deployer of the Services, the Customer is responsible for: (i) using the Services in accordance with the instructions for use provided by NuMind; (ii) implementing appropriate technical and organizational measures to ensure human oversight of Outputs; (iii) monitoring the operation of the Services and suspending use if a risk to safety or fundamental rights is identified; (iv) maintaining logs under the Customer's control for the minimum period required by applicable law, and in any event not less than six (6) months; (v) ensuring a sufficient level of AI literacy among Authorized Users and other persons dealing with the operation of the Services on the Customer's behalf; (vi) where required by the AI Act, informing natural persons that they are subject to the use of an AI system; (vii) where required by the AI Act, informing workers' representatives before deploying the Services in a context affecting working conditions; and (viii) promptly informing NuMind of any serious incident or malfunction identified during use of the Services, to the extent reasonably necessary to enable NuMind to fulfill its post-market monitoring obligations. #### 15.6 Provider Requalification If the Customer: (i) places the Services on the market or puts them into service under its own name or trademark; (ii) makes a substantial modification to the Services; or (iii) changes the intended purpose of the Services such that they become a high-risk AI system, the Customer shall be deemed the provider of the resulting AI system under Article 25 of the AI Act and shall assume all corresponding regulatory obligations. Such actions shall also constitute a material breach of the Agreement entitling NuMind to terminate the Agreement in accordance with Section 16.2. #### 15.7 Fundamental Rights Impact Assessment Where the Customer is required to conduct a fundamental rights impact assessment pursuant to Article 27 of the AI Act, NuMind shall, upon reasonable written request, provide the Customer with information necessary to conduct such assessment, to the extent such information is in NuMind's possession and not already available in the Documentation. The Customer shall treat such information as Confidential Information. The Customer remains solely responsible for the performance, methodology, and conclusions of such assessment. #### 15.8 Regulatory Updates Each party shall promptly notify the other of any changes in applicable law, including the AI Act, that materially affect the parties' respective obligations with respect to the Services. The parties shall cooperate in good faith to update the Agreement as necessary to reflect such changes. ### 16. Term, Renewal, Suspension and Termination #### 16.1 Term and Renewal Subject to terms of specific Order Form or any other specific agreement, the Agreement takes effect on the Effective Date and shall be renewed by tacit agreement for successive one (1) month periods, unless either party provides written notice of non-renewal prior to the expiry of the then-current term. #### 16.2 Termination In the event of a breach by either party of any of its obligations under the Agreement, the Agreement may be terminated at the fault of the defaulting party. Thus, in the event a party sends to the other notice of termination, by registered letter with acknowledgement of receipt, for failure to comply with one of its obligations under this Agreement: (i) if the breach may not be cured, the Agreement shall be immediately terminated by the non-defaulting party at the other party's fault following first presentation of said notice; (ii) if the breach may be cured, the defaulting Party shall have a period of fifteen (15) days as from the date of first presentation of said notice to definitively remedy to the breach or default. In this second case, if the breach or default is not definitively remedied within this period and the formal notice remains unsuccessful, the Agreement shall be terminated as of right at the defaulting party's fault. Termination is without prejudice to any damages to which the non-defaulting party may be entitled as a result of the breach or default by the defaulting party and to any recourse relating to the breach(es) found. The exercise of this right of termination does not exempt the defaulting party from fulfilling the obligations entered into until the termination takes effect, without prejudice to any recourse that the other Party may have. #### 16.3 Suspension NuMind may suspend the Customer's access to the Services immediately upon written notice if: (a) the Customer materially breaches any obligation under Section 4 or Section 5; (b) the Customer's use of the Services creates a material security, legal, or regulatory compliance risk for NuMind or third parties; or (c) any undisputed amount remains unpaid fifteen (15) calendar days after written notice from NuMind. Suspension does not relieve the Customer of its obligation to pay all Fees accrued and consumed prior to suspension. NuMind may reinstate access upon remediation of the grounds for suspension, at NuMind's reasonable discretion. #### 16.4 Effects of Termination 16.4.1 Upon expiry or termination of the Agreement for any reason: (i) all licenses and API access rights granted to the Customer terminate immediately; (ii) the Customer shall immediately cease all use of the Services and API Credentials; and (iii) all Fees for Tokens consumed up to the effective date of termination shall remain due and become immediately payable. 16.4.2 In accordance with Section 7.3, Customer Content and Outputs are deleted from NuMind's active production systems within fourteen (14) days of each API call. Residual copies may remain in encrypted backup or database history systems for up to thirty (30) days and are deleted through the ordinary backup rotation process. Any Customer Content held incidentally in connection with support services shall be deleted within thirty (30) calendar days after the Customer's written request or the effective termination date, unless retention is required by applicable law. 16.4.3 All amounts already paid to NuMind are retained by NuMind, and all amounts due for Tokens consumed shall be immediately payable. 16.4.4 The following Sections survive expiry or termination: Sections 1, 7.3, 7.4, 9 (Fees accrued prior to termination), 10, 11, 12, 13.4, 14, 15 (to the extent of obligations surviving by their nature), 16.4, and 19, together with any other provision that by its nature is intended to survive. ### 17. Modifications to the Terms and the Services 17.1 NuMind may modify these Terms at any time. In the event of any material modification, NuMind shall notify the Customer by email, or by any means agreed by the Parties where applicable, at least thirty (30) days before such modifications take effect. The Customer's continued access to or use of the Services following expiry of the notice period shall constitute the Customer's full acceptance of the modified Terms. 17.2 If the Customer objects to the proposed modifications, it shall notify NuMind in writing before the end of the notice period. In such event, the Customer: (a) shall not access or use any new features or functionalities introduced or modified after the effective date of the modifications; and (b) may terminate the Agreement in accordance with Section 16 before the modifications take effect. 17.3 NuMind may modify or discontinue the Services or any feature or functionality thereof at any time. Where a material change to the API or the Services would require the Customer to update its technical integration, NuMind shall use commercially reasonable efforts to provide advance notice using the notification channels specified in Section 19.6, with as much lead time as is reasonably practicable. ### 18. Governing Law and Jurisdiction The Agreement is governed by the laws of the State of Delaware, United States of America. In the event of a dispute, the parties shall seek to resolve it amicably in good faith. If no amicable resolution is reached within one (1) month after written notice of the dispute, the courts of the State of Delaware, United States of America, shall have exclusive jurisdiction. ### 19. General Provisions #### 19.1 Force Majeure Neither party shall be liable for any delay in or failure to perform its obligations under the Agreement caused by an event beyond its reasonable control, including acts of government, acts of God, war, terrorism, pandemic, major Internet or infrastructure outages, labor disputes, or natural disasters, provided such event is (i) beyond the reasonable control of the affected party, (ii) unforeseeable at the time of execution of the applicable Order Form, and (iii) unavoidable despite the exercise of all reasonable measures. The affected party shall notify the other party promptly of any force majeure event. If such an event continues for more than sixty (60) calendar days, either party may terminate the Agreement upon written notice, without prejudice to any amounts already due. #### 19.2 Subcontractors NuMind may use subcontractors and sub-processors to perform all or part of the Services, provided that NuMind remains responsible for the performance of the Services and for the acts and omissions of its subcontractors in connection with the Agreement. NuMind's use of sub-processors for the processing of personal data is governed by the DPA. #### 19.3 Assignment The Customer shall be solely responsible for the performance of the Agreement, and in particular shall refrain from assigning or transferring the rights defined in the Agreement. In addition, and to avoid all doubt, any changes which could occur in the person of NuMind, such as for example change of control, merger, scission, takeover, partial business transfer, assignment, transfer to a subsidiary, as well as any commercial or legal agreement with a third party, shall have no effect whatsoever on the existence and performance of the Agreement between NuMind and the Customer. #### 19.4 Publicity Unless the Customer notifies NuMind in writing at any time that it objects to being referenced, NuMind is authorized to identify the Customer as a customer and use the Customer's name, brand and logo on NuMind's website, customer lists, and standard sales and marketing materials, in each case in accordance with any brand guidelines made available by the Customer. The Customer may withdraw such permission on written notice, and NuMind shall cease such use upon such notice. In any event, these elements shall only be used in cooperation between the parties and in strict compliance with the Customer's image and reputation; the Customer retains full control of its image and may give NuMind any specific directions, agreements or refusal concerning the use of said elements within the framework of the present Agreement. #### 19.5 Open-Source and Third-Party Components The Services may incorporate open-source or third-party components subject to their own license terms. To the maximum extent permitted by applicable law, such components are provided on an “as is” basis. NuMind shall provide information on applicable open-source licenses upon reasonable written request. #### 19.6 Notices All notices under the Agreement shall be in writing. Routine operational notices may be sent by email to the contacts specified in the applicable Order Form. Formal notices relating to breach, termination, material claims, or registered mail obligations shall be sent by registered mail with acknowledgement of receipt (or an internationally recognized equivalent tracked delivery method) to the address defined above regarding NuMind and the address provided to the Customer regarding the latter. Either Party may update its notice details by written notice to the other party. #### 19.7 Entire Agreement The Agreement constitutes the entire agreement between the parties with respect to its subject matter and supersedes all prior proposals, negotiations, discussions, representations, and understandings relating to that subject matter, whether written or oral. #### 19.8 Waiver A failure or delay by either party in exercising any right or remedy under the Agreement shall not constitute a waiver of that right or remedy, nor shall it prevent or restrict any future exercise of that right or remedy. #### 19.9 Severability If any provision of the Agreement is held to be invalid, illegal, or unenforceable by a court of competent jurisdiction, such provision shall be deemed severed from the Agreement and the remaining provisions shall continue in full force and effect. The invalid provision shall be replaced by a valid provision that most closely reflects the parties' original intent. #### 19.10 Independent Contractors The parties are independent contractors. The Agreement does not create, and shall not be construed to create, a partnership, joint venture, employment, agency, or fiduciary relationship between the parties. #### 19.11 Electronic Signatures The Agreement, any Order Form, and any amendment may be executed by electronic signature or exchanged in electronic form, which shall have the same legal force and effect as original hand-written signatures. #### 19.12 Customer Purchasing Terms Any purchase order, procurement portal term, supplier onboarding document, or other customer-generated document shall not amend or supplement the Agreement. In the event of any contradiction between these Terms and any general terms and conditions of purchase or any other general or special terms and conditions of the Customer, these Terms shall prevail. #### 19.13 Language These Terms are written in English. In the event that these Terms are made available or translated in another language, the English version shall prevail in the event of any conflict or inconsistency between the different language versions. #### 19.14 Export Controls The Customer represents and warrants that: (a) it is not located in, incorporated in, or subject to the laws of a country that is subject to a U.S. government embargo or has been designated by the U.S. government as a terrorism-supporting country; (b) it is not listed on any U.S. government list of prohibited or restricted parties; and (c) it will not use the Services in violation of applicable export control or sanctions laws. # NuExtract Blog — Product Updates and Engineering Insights https://about.nuextract.ai/blog ## NuExtract Blog — Product Updates and Insights Product updates, release notes, and engineering deep-dives on NuExtract. Latest ### NuExtract3: The Reasoning Open-Source OCR & Structured Extraction LLM We introduce NuExtract3, a 4B open-source VLM specialized in document extraction. NuExtract3 unifies structured extraction (documents to JSON) and content extraction (OCR) into a single model. Trained via Reinforcement Learning to develop extraction-specific reasoning abilities (switchable on/off), it outperforms similarly sized models in both tasks — the new reference for open-source document extraction. May 19, 2026 July 16, 2025 #### NuExtract Platform: The New Information Extraction Today we release the NuExtract platform, a solution to extract high-quality structured information (JSON) from documents via API. This platform is powered by the recent NuExtract 2.0 PRO — the state-of-the-art LLM for information extraction. You can try it at nuextract.ai, or talk to us to get a private installation. July 16, 2025 #### NuExtract 2.0: Outclassing Frontier LLMs in Information Extraction We introduce NuExtract 2.0, the latest version of our LLM specialized in extracting structured information (document to JSON). NuExtract 2.0 brings vision, abstraction, and in-context learning abilities. We release open-source versions in the 2B-8B parameters range, and give access via API to our biggest model — NuExtract 2.0 PRO — which largely outperforms GPT-4.1 (+9 F-Score) and other frontier models. October 14, 2024 #### NuExtract 1.5 — Multilingual, Infinite context, still small, and better than GPT-4o! We introduce NuExtract 1.5, the new version of our foundation model for structured extraction. NuExtract 1.5 is multilingual, can handle arbitrarily long documents, and outperforms GPT-4o in English while being 500 times smaller. As usual, we release it under MIT license. June 24, 2024 #### NuExtract: A Foundation Model for Structured Extraction We introduce NuExtract, a lightweight text-to-JSON LLM. NuExtract allows to extract arbitrarily complex information from text and turns it into structured data. This model can be directly used in a zero-shot setting or fine-tuned to solve a specific extraction problem. As usual, we open-source it under MIT license for everyone to use. November 7, 2023 #### A Foundation Model for Entity Recognition Entity recognition is a widely used information extraction task, yet publicly available foundation models are not well suited for it. We leverage modern LLMs to create a small-yet-powerful foundation model for this task. This BERT-size model can be used to create custom entity recognizers with typically 5x less annotated data than before. This model is powering NuMind and we open-source it with an MIT license for everyone to use. Spread the word! August 25, 2023 #### Creating Task-Specific Foundation Models with GPT-4 There are two kinds of BERT-size NLP models in this world: general-purpose ones (a.k.a. foundation models), and highly specialized ones, trained on specific tasks and data. Neither kind is ideal to solve particular NLP problems on your data. We need to fill this specialization gap with task-specific foundation models, and propose a way to create them efficiently using LLMs. We apply this method to create a state-of-the-art domain-agnostic foundation model for Sentiment Analysis that we open source for everyone to use. June 21, 2023 #### What are Large Language Models? These past few months, thanks to ChatGPT and its siblings, we have been witnessing something historic. It seems that computers are finally able to understand our language, and are even able to speak back! These AIs are the latest iterations of large language models, also known as LLMs. But what exactly are these LLMs? How do they work? And how are they created? Let's dive into it. March 30, 2023 #### Seed Round Completed We are delighted to announce that we closed a seed funding round! We raised $3M, which we will use to continue building our NLP tool focused on text understanding. # NuExtract3: The Reasoning Open-Source OCR & Structured Extraction LLM — NuExtract Blog https://about.nuextract.ai/blog/nuextract-3-release Back to blog ## NuExtract3: The Reasoning Open-Source OCR & Structured Extraction LLM Alexandre Constantin Machine Learning Scientist Nathan Fradet Machine Learning Scientist Sören Dréano Machine Learning Engineer Etienne Bernard Co-Founder & CEO May 19, 2026 We introduce NuExtract3, a 4B open-source VLM specialized in document extraction. NuExtract3 unifies structured extraction (documents to JSON) and content extraction (OCR) into a single model. NuExtract3 is trained via Reinforcement Learning to develop extraction-specific reasoning abilities, which can be switched on and off on demand. We find that NuExtract3 outperforms similarly sized models in both structured and content extraction, making it the new reference model for open-source document extraction. NuExtract3 unifies structured extraction (JSON) and content extraction (Markdown/OCR) into a single model, powering business process automation and AI agents/assistants. ### Quick Links - 🖥️ NuExtract Platform — to use NuExtract - 🤗 NuExtract3 on HuggingFace - 📁 GitHub Repository — inference & fine-tuning scripts - 🗣️ Discord — support & updates ### TL;DR - We release NuExtract3 , an open-source 4B VLM specialized in document extraction. - NuExtract3 unifies: - Structured extraction (document to JSON via a template) - Content extraction (document to Markdown, a.k.a. OCR) - NuExtract3 outperforms similarly sized models in both structured and content extraction. - Key features: - Reasoning abilities you can turn on and off - Support for freeform instructions - Support for in-context examples - 20 structured extraction field types - NuExtract3 is based on Qwen3.5-4B, and released under Apache 2.0 license. ### A Unified Document Extractor Modern document extraction is split into two main tasks: Structured extraction and Content extraction . Structured extraction involves extracting specific information from a document and returning it in a computer-readable format, typically a JSON file. The extraction is defined by a schema (a.k.a. template): Structured extraction task. A machine-readable structured output (JSON) is extracted from a document according to a schema/template. This task automates data entry for banks, insurances, and healthcare organizations. Such extraction is widely used by banks, insurance companies, and healthcare organizations to capture key information (names, addresses, amounts, etc.) from incoming documents (invoices, claims, paystubs, etc.) and enter it into their systems. In other words, structured extraction automates data entry . Content extraction (a.k.a. OCR) is conceptually simpler. It involves extracting the document's full content and meaning into a text-based format, typically a Markdown file: Content extraction task (a.k.a. OCR). A full-content Markdown is extracted from a document. This extraction makes documents accessible to AI assistants/agents. Such extraction is mainly used to pre-process enterprise documents (PDFs, scans, etc.) so they are accessible to LLMs, typically in a RAG setup. In other words, content extraction makes enterprise documents AI-ready . Note that we use the terms "content extraction" and "OCR" (Optical Character Recognition) interchangeably, but content extraction also applies to digitally native documents, not just scanned ones. At NuMind, we've been building specialized document extraction models for the past two years. We first focused on structured extraction with our NuExtract line of models, which saw broad adoption ( 2M+ downloads ). We then leveraged our experience building VLMs to tackle OCR and created NuMarkdown-Thinking-8B , which became our most popular model (1.5M+ downloads). We now want to pursue both directions, but with a unique line of models… Indeed, while structured and content extraction have different applications, they rely on the same core capability: document understanding. So why not unite both tasks in a single model? As it turns out, this approach works well. We find that a unified model is more robust and performs better than models trained separately. It also simplifies deployment when you need to tackle both tasks. We therefore decided to create NuExtract3, the first unified OCR & structured extraction model. ### A Reasoning Extractor VLMs of all kinds, even the largest ones, struggle with complex documents, such as those containing handmade tables: A document that is typically challenging for LLMs. The table is split in two, with no repeated headers. A row is split into three sub-rows. Content in one cell overlaps with adjacent cells. Typical issues arise when multiple tables sit side by side, when rows or columns split, or when cell content overflows into nearby cells. To tackle these issues, we found that a reasoning approach was effective. Before providing the actual result, the model thinks "out loud" about the document, moving from the general (e.g. sections) to the specific (e.g. header names) while anticipating potential pitfalls . Here is an example of reasoning from our earlier model NuMarkdown-Thinking-8B, the first reasoning OCR model: Thinking trace from NuMarkdown-Thinking-8B. The model thinks from the general to the specific while anticipating potential pitfalls. This kind of reasoning is effective at resolving document understanding issues, making NuMarkdown competitive with models that are much larger. We decided to bring this reasoning ability to NuExtract, both for OCR and structured extraction . Reasoning, however, is not "free". It requires generating thinking tokens for each extraction, which increases cost and latency. This is a major issue with generalist models, which often generate ten times more thinking tokens than output tokens, multiplying cost and latency by the same factor. To address this, we trained NuExtract3 to use roughly the same number of thinking tokens as output tokens . We find this to be a sweet spot between extraction quality, cost, and latency. We also made it possible to turn reasoning on and off . ### Making NuExtract3 Building a task-specialized model starts with a dataset. This dataset must be both diverse and challenging. For diversity, we use real-world documents from Fine-PDF , which we automatically annotate through a range of processes, including LLMs as annotators and judges, iterative corrections, and programmatic filtering. For difficulty, we generate complex synthetic documents, which have the advantage of being perfectly annotated. The next step is to pick a base model. We choose the open-source generalist model Qwen3.5-4B. This model has impressive general capabilities for its size, which gives us a strong starting point. We then train the base model on our dataset in two phases. First, we use supervised learning via next-token prediction. Second, we use reinforcement learning. While the first phase is already effective at specializing the base model, the second is necessary for better template adherence and stronger reasoning abilities. Here is a summary of the process: Training procedure for NuExtract3. We create an extraction training set and use it to fine-tune Qwen3.5 4B via supervised fine-tuning and reinforcement learning. ### Structured Extraction Results Let's now look at the quality of NuExtract3's extractions, starting with structured extraction (JSON output). To do so, we use our structured extraction benchmark, which includes about 600 challenging extractions across 15 different problems. We predict extraction trees in a zero-shot setting and compare the extracted leaf values with their ground-truth values. We use the EXTRA metric, which is essentially leaf accuracy (we are currently writing a paper about this benchmark & metric). Here is a comparison of NuExtract3 with the best models of a similar size: Model performance comparison for the structured extraction tasks (doc to JSON). NuExtract3 outperforms all similarly sized models. We can see that NuExtract3 outperforms generalist models substantially , beating Gemma 4 by more than 10 points. One aspect that does not work as well for such generalist models is reasoning. While it is beneficial for Gemma, it is detrimental for Qwen, GLM, and Ministral, where the thinking trace often loops. We specifically train NuExtract3's thinking via reinforcement learning, which fixes these issues. ### Content Extraction (OCR) Results Let's now evaluate NuExtract3's OCR capabilities. We want to understand how effective NuExtract3 is at preprocessing documents for an LLM to use later on . While there are plenty of OCR benchmarks out there, we found they do not properly measure the capabilities we care about. They either focus on character recognition, like OCRBench v2 , on preserving document layout, like OmniDocBench , or, like olmOCR-Bench , on preserving document semantics via hand-crafted programmatic metrics, which can lead to suspicious results, like Qwen3.5 2B scoring higher than Gemini 3.1 Pro. We thus decided to figure out another way to test these models. As a first baseline, we use a frontier LLM as a judge (Gemini 3.1 Pro with maximum thinking) to compare models. We provide a pair of Markdown outputs and ask the judge to tell which one is best via this naïve prompt: Prompt to judge two Markdown outputs. We use 150 complex documents (mostly weird tables) and compare extracted markdowns between NuExtract3 and models of similar size, both generalist and specialized. Here is the win rate of these models against NuExtract3: Model performance comparison for the content extraction tasks (OCR) via an LLM judge. NuExtract3 outperforms all similarly sized models. Like in the case of structured extraction, NuExtract3 largely outperforms generalist models at content extraction , showing once again the usefulness of specialized models. NuExtract3 also scores higher than specialist models, although LightOnOCR 2 and Chandra OCR 2 score pretty high for their size. This "OCR-battle benchmark" gives a first view of NuExtract3 capabilities, but it has many flaws: it only includes 150 documents, the LLM judge is not perfect, and, importantly, we baked-in handcrafted features to define what a "good Markdown extraction" is, which is human-biased. For example, this benchmark could be too sensitive to styling as opposed to extracting a useful Markdown. We believe there is a much better way to test OCR abilities… The "obvious" thing to do would be to evaluate these OCR models based on how well an LLM can use their Markdown outputs to answer questions , since that is their purpose. However, this is not easy: you would need a large set of high-quality questions, one LLM to answer them, and another LLM to evaluate the answers. One way to simplify this could be to ask simple questions with answers that can be verified programmatically, removing the need for an evaluator. As it turns out, this is exactly what structured extraction is, and we already have a benchmark for that. We thus decided to repurpose our structured extraction benchmark to measure OCR capabilities . For each model, we extract Markdown files for all 600 documents in the benchmark, then use a standard LLM (Qwen3.6 27B here) to extract structured information from the Markdowns. This provides us with around 100k extracted leaf values to test the model on. We then compare these values to the ground truth using the EXTRA metric (essentially, leaf accuracy). This testing approach is free of styling bias: it directly measures how useful the results are for an AI . We believe that this is the current best way to test OCR models. Here is what we obtain: Content extraction (OCR) performance. We extract Markdown files from 600 documents, and measure the ability that a standard LLM (Qwen3.6 27B) has to extract structured information from these Markdowns. NuExtract3 outperforms all similarly sized models. We can see that NuExtract outperforms both generalist and specialist models while using an average of only 338 thinking tokens. Generalist models compete with the specialized ones, but at the cost of a high number of thinking tokens (an average of 6,552 for Qwen and 1,973 for GLM). We can see the benefits of reasoning for this task. NuExtract would likely benefit from generating more thinking tokens before answering; we were probably too aggressive in penalizing thinking tokens in this version. It is interesting that generalist models perform much better here than in the OCR battles. We believe this is partly due to styling bias in the judge LLM. In the end, we do not care about reproducing the exact document layout. We only care about preserving the information and making it easy to process. ### New Field Types (Structured Extraction) NuExtract 2.0 introduced field types in the template: Example of a NuExtract 2.0 template and a compatible extraction output. There were 7 possible types, and notably "verbatim-string" , which tells the model to extract without any reformulation, reducing hallucinations: Type Description Example "verbatim-string" copy-pasted string from input document "John" "string" any string "USD" "integer" a whole number 8 "number" a whole number or a decimal number 1.39 "boolean" true or false true ["x_1","x_2",…] Choice between strings "x_1" , "x_2" , etc. "Large" "date-time" date and/or time in ISO 8601 , as a string "2023-03-25T04:32:17" NuExtract3 introduces 14 new types: Type Description Example "date" date in ISO 8601 "2023-03-25" "time" time in ISO 8601 "04:32:17" "duration" duration in ISO 8601 (PnYnMnDTnHnMnS) "P2Y1M3D" "country" country in ISO 3166-1 "FR" "currency" currency in ISO 4217 "EUR" "language" language in ISO 639-3 "eng" "language-tag" IETF language tag "en-US" "url" IRI in RFC 3987 "http://www.example.com/" "email-address" email address in RFC 5322 and RFC 6531 "用户@例子.公司" "phone-number" phone number as a sequence of digits or in ITU E.164 "+14155552671" "iban" IBAN in ISO 13616-1 "DE89370400440532013000" "bic" BIC in ISO 9362 "BNPAFRPPXXX" "unit-code" UCUM unit code "m" "region:XX" Regions of a given country. "XX" can be "US" , "FR" , etc. "MA" These types allow for more precise control over how extracted values are represented, which reduces the need for post-processing. ### Model Instructions With NuExtract 2.0, the only way to improve structured extraction performance was to add in-context examples or do some "template engineering", which consists of adding or modifying field names. This sometimes resulted in long field names like "card_access_number//The one on the bottom right" , which uses additional tokens. For NuExtract3, we added the ability to provide additional instructions to the template. For example, let's say you want to extract the "card access number" from this ID card: Extracting the card access number from a French national ID card. Instead of putting additional information in the field name, you can add it in the instructions, such as "The card access number is 6 digits and generally located at the bottom right of the card." , which helps distinguish this number from the document number. Model instructions were a highly requested feature and, while we are not totally satisfied with how they are followed by the model, we find using instructions extremely useful . ### Long live NuExtract 3! That's it for this release. NuExtract3 shows that a specialized model can lead its category in both structured extraction and OCR. We hope that this model will be useful for your projects. As usual, please share feedback on what we should prioritize for NuExtract4 (model confidence? bounding boxes? instructions for content extraction?). # NuExtract Platform: The New Information Extraction — NuExtract Blog https://about.nuextract.ai/blog/nuextract-platform-the-new-information-extraction Back to blog ## NuExtract Platform: The New Information Extraction Francois Laroche Senior Software Engineer Charles Mativat Senior Software Engineer Olga Chuchuk Software/ML Engineer Samuel Bernard Co-Founder & CTO Etienne Bernard Co-Founder & CEO July 16, 2025 Today we release the NuExtract platform, a solution to extract high-quality structured information (JSON) from documents via API. This platform is powered by the recent NuExtract 2.0 PRO — the state-of-the-art LLM for information extraction. You can try it at nuextract.ai, or talk to us to get a private installation. This post is about the NuExtract Platform — check here for the sister post about NuExtract 2.0 . ### TLDR - The NuExtract platform allows to extract high-quality structured information (JSON) from documents - Documents can be texts, PDFs, spreadsheets, scans, etc. in any language - It is powered by NuExtract 2.0 PRO — our specialized LLM outperforming frontier LLMs - You can create tasks (template + examples) & test the model via the web interface - You can extract information at scale via API ($5 per million tokens) - Can be deployed privately for full data confidentiality - Quick Links: - 🖥️ NuExtract Platform - 📖 Platform Documentation - API Reference & Python SDK - 🗣️ Discord community - 📹 Video Tutorial ### Information Extraction Information extraction — sometimes called structured extraction — is the task of extracting information from an unstructured document (email, invoice, form, contract, and so on) into a structured output for a computer to use . For example, let's say that you need to verify online users. You would need to transform scans of their IDs into structured data: Structured extraction from a scanned ID. The output is in JSON format. Structured extraction from a scanned ID. The output is in JSON format. The above JSON output can easily be handled by a computer. Similarly, you might need to extract quantities/prices from invoices or receipts: Structured extraction from a receipt of payment in Chinese. Structured extraction from a receipt of payment in Chinese. Or you might need to classify/extract information from technical documents such as this plan: Structured extraction from a floor plan. Structured extraction from a floor plan. Information extraction is not just about processing scanned documents. More often than not, documents are just a regular PDFs, like this contract: Legal contract extraction. Legal contract extraction. It can also be about classifying or extracting information from raw text documents, such as reports, emails, or customer messages: Text classification & text extraction. Text classification & text extraction. If the input is a document and the output is a JSON, this is an information-extraction task . Companies are flooded with unstructured documents; there are needs for information extraction about everywhere. Here are the main use cases that we encounter at NuMind, organized by industry: Main structured-extraction use cases organized by industry. Main structured-extraction use cases organized by industry. If you work in such industry, chances are that you have information extraction needs as well ! ### Why the NuExtract Platform? The field of information extraction has a long history. From decades ago, and up until recently, extraction was tackled via heuristics, regex-like rules, shallow ML methods, traditional CR preprocessing, and a lot of human effort. These methods limited information extraction to simple tasks, with low-variability documents . Large Language Models (LLMs) are changing the deal. Thanks to their language understanding, world knowledge, and ability to generate complex outputs, they solve extraction problems that were previously out-of-reach. Furthermore, as they continue improving, LLMs hold the promise of "solving" information extraction entirely , which means being able to perform any extraction task perfectly while only requiring minimal human input to define the task. Diagram showing information extraction moving from rules and shallow machine learning to large language models capable of complex extraction tasks This promise is not satisfied yet — LLMs still make plenty of extraction mistakes — but we found a path to get there: We discovered about a year ago that it was possible to create specialized LLMs that were much better at extracting information than generalist LLMs . From this research, we created NuExtract, a line of LLMs specialized in extracting information. One interesting thing is that NuExtract models hallucinate less , as we managed to teach them to say "I don't know" ( null value, in JSON speak) when the requested information is not present in the document. Making of NuExtract 2.0. Making of NuExtract 2.0. At first, NuExtract models were small — only a few billion parameters — and limited to text documents. We recently moved on to bigger models which can also process PDFs & scans via a vision module . The nice surprise is that performance gains that we see on small models are still present on big models! Our latest and biggest model to date, NuExtract 2.0 PRO, is simply outclassing non-reasoning frontier LLMs : 0-shot performance of NuExtract 2.0 PRO compared to non-reasoning frontier models on the extraction benchmark (text and image documents). NuExtract outperforms GPT-4.1 by over 9 F-Score points. 0-shot performance of NuExtract 2.0 PRO compared to non-reasoning frontier models on the extraction benchmark (text and image documents). NuExtract outperforms GPT-4.1 by over 9 F-Score points. Even more surprising, NuExtract 2.0 PRO is also surpassing reasoning frontier models, while being faster and at least 10x cheaper to use : 0-shot performance of NuExtract 2.0 PRO compared to reasoning frontier models on the extraction benchmark (text and image documents). NuExtract outperforms o3 by 3 F-Score points. 0-shot performance of NuExtract 2.0 PRO compared to reasoning frontier models on the extraction benchmark (text and image documents). NuExtract outperforms o3 by 3 F-Score points. These results motivated us to create the NuExtract platform , mostly to provide API access to NuExtract 2.0 PRO. The NuExtract platform ended-up being much more than "just" providing API access to a model. For example, we included a pre-processing step to handle various document formats (PDFs, spreadsheets, scans) , a post-processing to make sure the output is a valid JSON, and, importantly, a graphic interface to easily define extraction tasks and test the model . Here is what this interface looks like: NuExtract platform graphic interface. Create a template, test the model, and extract via API. NuExtract platform graphic interface. Create a template, test the model, and extract via API. Note that this platform, like the model NuExtract 2.0 PRO, is multilingual . Another important thing is that the NuExtract platform can be deployed privately , which is a requirement for companies having data privacy/confidentiality constraints. This is in part due to the fact that NuExtract 2.0 PRO, while being the biggest of the NuExtract models, still fits on one H100 GPU , making it practical for private use. Overall, there are four main reasons for using the NuExtract Platform over alternative solutions: - Highest extraction quality , thanks to NuExtract 2.0 PRO - Reasonable price compared to frontier LLMs. - Ease of use (task creation & testing via the web platform, task-specific API endpoint) - Possibility of private use to ensure data confidentiality ### How to Use It You can use the entire platform via API if you want, and directly make extraction calls . However, the friendlier way to use the platform is to: - Define your extraction task in the user interface, which means: - Creating a template - Adding extraction examples to teach NuExtract - Testing the model in the playground - Deploy to production via an API endpoint specific to your extraction task. Let's look at these steps further (and you can check the user guide for more details). #### Creating a Project The very first step is to create a "project". In the platform, a project corresponds to one specific extraction task , and each project has its own API endpoint to extract information from documents. You can choose to start a project from scratch, or duplicate an existing "reference project": Start a new project by creating one from scratch or by duplicating an existing reference project. Start a new project by creating one from scratch or by duplicating an existing reference project. Each project has four tabs: - Workspace — to create the template and play with NuExtract - Example Set — to improve extraction by adding teaching examples - API — to deploy into production - Settings — to control things like model temperature and rasterization resolution of PDF documents. #### Template The next step is to create a template for your extraction task in the Workspace. The template defines what to extract and how the output should be structured — it is a hard constraint on the output. Importantly, the returned output always matches its corresponding template . Here is an example of template and compatible extraction output: A template and compatible extraction output. A template and compatible extraction output. You can see named fields such as "first name" indicating what to extract, type specifications such as "verbatim-string" indicating types/formats that the extracted values should have, and constructors {$...$} (object) and [$...$] (set) defining the output structure. To create such a template, you can provide a description of the task , such as "extract minimal information from a CV", and press the magic wand 🪄 to obtain a valid NuExtract template: Creating a template from a task description. Creating a template from a task description. You can then modify the result to be exactly what you want by looking at the template format in the documentation. Instead of a description, you can also provide a JSON schema, Pydantic code, or even a document — pressing the magic wand will turn whatever you provide into a NuExtract template. #### Teaching Examples Templates alone can be ambiguous; we sometimes gain to give NuExtract examples of our task . This can be done in the "Example Set" tab by providing input→output examples of correct extractions, such as: Input→output example of the extraction task to teach NuExtract. Input→output example of the extraction task to teach NuExtract. NuExtract learns from such example to perform the task better . Even a unique example can improve performance substantially. It is generally a good idea to provide examples for which the model struggles . These examples are added to the prompt of NuExtract (a.k.a. in-context learning), which means that, at the moment, the number of examples is limited by the context size of the model. We are working on a solution to allow for an arbitrary large number of examples. #### Playground At any point during the definition of your task, you can test the model in the playground , either by typing/pasting text, or by uploading a document: Playground result from the template above. Playground result from the template above. The extraction here seems correct, and you can see that some fields have null values. null is the way NuExtract expresses that it could not find or infer the requested information . Knowing when to return null is a strength of NuExtract. The goal of the playground is generally to try to find extraction errors. When you managed to find such error, you can add the corresponding document and corrected extraction to the teaching examples , in order to correct the model. Note that you can create multiple "playpods" to keep track of performance as you modify the template and teaching examples. ### Extracting via API Once you are happy with how NuExtract behaves for your task, it is time to put it in production! To do so, there is one extraction API endpoint for each project: https://nuextract.ai/api/projects/{projectId}/extract You provide a text or a file, and it returns the extracted information according to the task defined in the project. To use this endpoint, you need to create an API key and replace {projectId} by the project ID found in the API tab of the project. Let's test it on a minimal text document: API_KEY="*your_api_key_here*"; \ PROJECT_ID="87f22ce1-5c1d-4fa9-b2f1-9b594060845f"; \ curl "https://nuextract.ai/api/projects/$PROJECT_ID/extract" \ -X "POST" \ -H "Authorization: Bearer $API_KEY" \ -H "content-type: text/plain" \ -d "Alice began attending Wonderland Academy on July 4, 1862." The result is: {"result": { "First name":"Alice", "Last name":null, "Skills":[], "Education":[ { "School":"Wonderland Academy", "Start date":"1862-07-04", "End date":null } ] }, "completionTokens":52, "promptTokens":237, "totalTokens":289, "logprobs":-0.13810446072918126 } We can see that the extraction is correct and that that null has been used to represent missing information. We can also see the number of input and output tokens, and the total log probabilities of output tokens, which can help figuring out the confidence of the model in its extraction . Similarly you can try this endpoint on a file document: API_KEY="*your_api_key_here*"; \ PROJECT_ID="87f22ce1-5c1d-4fa9-b2f1-9b594060845f"; \ curl "https://nuextract.ai/api/projects/$PROJECT_ID/extract" \ -X "POST" \ -H "Authorization: Bearer $API_KEY" \ -H "content-type: application/octet-stream" \ --data-binary @file_name.ext And you can also use the Python SDK to perform such extractions: from numind import NuMind from pathlib import Path project_id="87f22ce1-5c1d-4fa9-b2f1-9b594060845f" client = NuMind(api_key=api_key) file_path = Path("document.odt") with file_path.open("rb") as file: input_file = file.read() output_schema = client.post_api_projects_projectid_extract(project_id, input_file) You can find more information about the API in the API Reference and about the Python SDK in the SDK documentation . ### Pricing This platform follows a simple pricing model: - Everything done in the user interface is free - Using the extraction API costs $5 per million tokens - Everything else done via API is free Note that, in this pricing, we are mixing input tokens (template, examples, and input document) and output tokens (tokens of the generated JSON). Generally, the majority of tokens originate from input documents . To get an estimate of what this means for your documents, 1 word is about 1.3 tokens on average in English language, and, for image documents, one token corresponds to a patch of 28x28 pixels. Here is an estimation of input document prices: API price with NuExtract 2.0 PRO. API price with NuExtract 2.0 PRO. We are working on including a smaller model, priced under the $1 per million tokens bar. Note that if you need to process more than a few million pages a year, it might be worth considering a private NuExtract platform to reduce inference price (e.g. by batching documents or by using a fine-tuned model). Talk to us to know more about it. ### What Happens with Your Data? Now, let's talk a bit about the data you send to the NuExtract platform. This is an important topic since input documents might contain private/confidential information about persons and organizations. In a nutshell, we only keep what is needed for the platform to function , which means: - The current template and examples defining the task. - The current documents in the playground. Production documents and their extracted information are deleted in a maximum of two weeks after being processed . Also, importantly, we do not send anything to a third party . Documents are processed by our models on servers that we control. We do not send documents to external APIs or anything like that. Finally, we do not train models on any document sent to the platform. Now, we know these guarantees are not enough for everyone, which is why we also offer to host the NuExtract platform privately: on your private cloud, on your premises, or on a dedicated instance that we host for you. If this interests you — and until we make this private platform self served — you will have to talk to us about it 🙂. ### Limitations & What's Next NuExtract 2.0 PRO is the best at extracting information… but not perfect (yet) by any means! Also, the NuExtract platform is in its infancy, a lot to improve! Let's look at some limitations of both the model and the platform, and what we plan to address them. - Limited document size This is probably the biggest limitation of this platform. Because of the 32k-token context window of NuExtract 2.0 PRO, you are limited to about 60 pages of text, or 20 pages of images, which is not enough for some applications. We are working on a solution to fix this problem entirely. In the meantime, you can try to split the input document and merge information afterward. - Limited number of teaching examples Since teaching examples are included in the prompt, their size and number are also constrained by the 32k context size of the model. We figure out ways to bypass this limit, but it will have to wait a bit more before being released. - Lack of extra instructions The last main limitation is probably the inability to express subtleties about your task that are not easily expressed in the templates, and which would require to many examples for NuExtract to "get it". This should be relatively easy to fix. In the meantime, one trick to bypass this limitation is to add "feature fields" to guide the model. For example, to classify a resume as relevant or not, you might include fields like "Has candidate a business degree?". Besides these obvious limitations, there are plenty of new features waiting to be implemented. For example: - A proper way to measure and display model uncertainty - A way to visualize where the information was extracted from in the document - Better JSON visualization and edition (for creating examples) - A production monitoring interface - And, of course, even better & cheaper models We are working on all of these, and we need all the feedback we can get to debug, prioritize, and make design choices. Do not hesitate to let us know what you think 🙂. ### Extract baby, Extract! That's it for this post! We are thrilled to be working on this project and hope that it will useful to many of you, give it a try! 🚀 # NuExtract 2.0: Outclassing Frontier LLMs in Information Extraction — NuExtract Blog https://about.nuextract.ai/blog/outclassing-frontier-llms-nuextract-2-0 Back to blog ## NuExtract 2.0: Outclassing Frontier LLMs in Information Extraction Alexandre Constantin Machine Learning Scientist Liam Cripwell Machine Learning Scientist Nathan Fradet Machine Learning Scientist Sören Dréano Machine Learning Engineer Etienne Bernard Co-Founder & CEO July 16, 2025 We introduce NuExtract 2.0, the latest version of our LLM specialized in extracting structured information (document to JSON). NuExtract 2.0 brings vision, abstraction, and in-context learning abilities. We release open-source versions in the 2B-8B parameter range, and give access via API to our biggest model — NuExtract 2.0 PRO — which largely outperforms GPT-4.1 (+9 F-Score) and other frontier models. ### Quick Links - 🖥️ NuExtract Platform — to use NuExtract - 🤗 NuExtract 2.0 Models on HuggingFace - 📁 GitHub Repository — inference & fine-tuning scripts - 🗣️ Discord — support & updates ### NuExtract Goes On! NuExtract is an LLM specialized in extracting structured information from documents : it returns a JSON output from a document and a template (a.k.a. schema). We released NuExtract 1.0 about a year ago, NuExtract 1.5 about 6 months ago, and have been pleased to see their popularity (several million downloads) and performance (similar to GPT-4o while much smaller). Nevertheless, NuExtract 1.0 and 1.5 had important limitations: they could only perform pure extraction (copy-paste of the input) on text documents. We thus decided to add 3 key features to NuExtract 2.0: - Vision : to process documents as images, allowing extraction from scanned documents, PDFs, Excel files, etc. - Abstraction : to perform classification, reformulation, formatting, etc. - In-context learning : to customize the model by adding examples in the prompt. On top of that, we gave a strong push on the performance side of things — better datasets, better training procedures, and better base models. Creation procedure of NuExtract 2.0. We are releasing 3 open-source versions of NuExtract 2.0: - NuExtract 2.0 2B (Qwen 2.0 VL 2B base, MIT license) - NuExtract 2.0 4B (Qwen 2.5 VL 3B base, research license) - NuExtract 2.0 8B (Qwen 2.5 VL 7B base, MIT license) All have a 32k token context. These models perform impressively well. But what if we go bigger? We trained a substantially bigger model — NuExtract 2.0 PRO — and were stunned by the results: the model simply outclasses frontier LLMs like GPT-4.1 and Claude 4 Opus. 0-shot performance of NuExtract 2.0 PRO compared to frontier LLMs on the extraction benchmark (text and image documents). NuExtract outperforms GPT-4.1 by over 9 F-Score points. Driven by these results, we created a platform entirely dedicated to NuExtract 2.0 PRO — nuextract.ai — where you can define extraction tasks and use NuExtract via API. ### Giving Eyes to NuExtract Documents are not just text — they can be formatted (PDFs, spreadsheets) or scanned. The traditional approach captures their raw text via OCR, which loses formatting information: titles merge with paragraphs, tables lose structure, diagrams disappear. Structured extraction from a scanned ID card. A recent way to solve this is to directly extract from images with a Vision Language Model (VLM) . VLMs have made enough progress that traditional OCR is now largely unnecessary . We thus made NuExtract 2.0 a VLM, using Qwen 2.5 VL and Qwen 2.0 VL as base models. Images are tokenized by patches of 28×28 pixels and embedded via a vision module. The unified sequence of token embeddings is then processed by a regular transformer. There is no hard limit on image size: NuExtract 2.0 can process arbitrarily large (or high-resolution) images . We find the resulting model can correctly extract information from all kinds of image documents — receipts, floor plans, multi-page PDFs, even handwritten prescriptions (with fine-tuning for difficult cases). Structured extraction from a receipt. Structured extraction from a floor plan. Structured extraction from a multi-page PDF document. Prescription extraction — the final OCR frontier. Importantly, the multimodal transition does not affect performance on text documents, so we only release a multimodal model . ### Abstraction & New Template Format NuExtract 1.0 and 1.5 were pure extractors: they copy-pasted text from the input. Pure extraction is a common use case, but you sometimes need to go beyond it: - Reformat an output — dates, country codes, numbers. - Deduce an answer — country from capital, number of rooms from a floor plan. - Predict a class — sentiment, job mode (remote/on-site/hybrid). - Generate new text — translation, summarization. To get the best of both worlds, we updated the template format to include type specifications . In particular, the "verbatim-string" type tells the model to perform a pure extraction, while "string" allows free generation. We also added "choice" (enum), date, and number types. null now represents missing information. Template field types. Template constructors. Example of a NuExtract 2.0 template and a compatible extraction output. This is a minimalist template format, designed to be easy to read and write, both for humans and LLMs , while being precise about what to extract. ### In-Context Learning Templates resolve a lot, but not everything: in {"address": "string"} , the format is unspecified. Providing examples in the prompt — In-Context Learning (ICL) — is the natural answer. We taught NuExtract to perform ICL by adding examples-in-prompt to the training set, both for text and images. Minimalist training example to teach NuExtract to perform in-context learning. Even 3 examples can substantially improve performance : NuExtract 2.0 PRO gains 6 F-Score points, which is a lot at this level of performance. In-context learning abilities of NuExtract 2.0 on the extraction benchmark (text + image). ICL is a great way to provide a light but efficient fine-tuning to NuExtract. ### Performance We use our extraction benchmark composed of 1000+ extraction examples grouped into 21 extraction problems. Documents are text or images, and span multiple languages. Open-source models — transforming generic models into extraction specialists massively increases performance. NuExtract 2.0 8B reaches 73 F-Score (a bit better than non-reasoning frontier models). These small specialized language models are well suited for high-volume applications with resource-limited infrastructure. 0-shot performance of open-source NuExtract 2.0 models compared to their base models. NuExtract 2.0 PRO vs frontier models — NuExtract 2.0 PRO largely outperforms frontier models, with a +9 F-Score margin over GPT-4.1. vs reasoning models — NuExtract 2.0 PRO is also ahead of frontier reasoning models: +5 F-Score over reasoning Claude 4 Opus, +2 F-Score over Gemini 2.5 PRO. And it costs at least 10× less to use on extraction tasks. 0-shot performance of NuExtract 2.0 PRO compared to reasoning frontier models. NuExtract outperforms o3 by 3 F-Score points. Precision vs recall — for structured extraction, precision is king : you prefer not having information rather than polluting your database with wrong information. NuExtract 2.0 PRO has a higher precision than recall, in part because we specifically teach it to say "I don't know" ( null ) when information is missing. Recall (left bars) and precision (right bars) results on the extraction benchmark. ### Failure Points - Long documents — context size of 32k tokens (~60 pages of text or 20 pages of images) is the most obvious shortcoming. All LLMs also tend to have issues with long lists. - Laziness — on complex templates with long lists and many missing properties, NuExtract may return very little. Workaround: simpler templates or split documents. - Looping — rarely, NuExtract repeats the same element. This mainly happens on low-resolution images with non-Latin characters. The platform detects and corrects this. - Invalid JSONs — very rare (e.g. 07 instead of 7 ). The platform guarantees valid JSON output. Off-platform, use a library like jsonrepair . ### Conclusion & Next Steps NuExtract 2.0 is the highest-performing extraction LLM , usable for pretty much any extraction task, on any kind of document, in any language. The road is not over: better uncertainty estimates, longer documents, and reasoning are on the roadmap. In the meantime, we hope you'll make good use of this model. Feedback always welcome 😊 # NuExtract 1.5 — Multilingual, Infinite context, still small, and better than GPT-4o! — NuExtract Blog https://about.nuextract.ai/blog/nuextract-1-5-multilingual-infinite-context Back to blog ## NuExtract 1.5 — Multilingual, Infinite context, still small, and better than GPT-4o! Liam Cripwell Machine Learning Scientist Alexandre Constantin Machine Learning Scientist Etienne Bernard Co-Founder & CEO October 14, 2024 We introduce NuExtract 1.5, the new version of our foundation model for structured extraction. NuExtract 1.5 is multilingual, can handle arbitrarily long documents, and outperforms GPT-4o in English while being 500 times smaller. As usual, we release it under MIT license. - This post is a continuation of the original NuExtract post . - You can download NuExtract 1.5 here and try it here . ### Why NuExtract? NuExtract is a family of small open-source models that do only one thing: they extract information from documents and return a structured output (JSON) . Because they only do this one thing, they are very good at it — NuExtract 1.5 (3.8B parameters) is better than GPT-4o on our English zero-shot benchmark while being 500× smaller . Moreover, fine-tuning NuExtract gives you performance hard to reach via prompting, even for a frontier LLM. Such a small open-source model has two main advantages: - You can use it privately, without sharing your data. - You can fine-tune it on input-output examples to make it excel at a specific task. For repetitive tasks needing high performance, or for sensitive data, NuExtract is the way to go. ### NuExtract 1.5 in a Nutshell The two main requests we received after NuExtract 1.0 were the ability to process long documents and to handle non-English documents . In a nutshell, we created a multilingual dataset and trained the latest open-source LLMs on it . We then added a "continuation" functionality, which gives NuExtract an infinite context size with a bounded memory footprint . ### Multilingual Abilities We need both a multilingual dataset and a multilingual base model. Phi-3.5 mini recently made a lot of progress on that side, handling 23 languages. We chose Phi-3.5 mini as the base of NuExtract 1.5. For the training dataset we again take raw documents from the C4 dataset : 50% English, 50% other languages (mainly French, German, Spanish, Italian, Portuguese), with longer documents than for the original NuExtract. For half the documents we use an English template (irrespective of the document language), so users can write a unique template in English when processing multilingual documents. The other half uses templates in the document's own language. The dataset stays purely extractive — copy-paste, no generation. Training example with a French document and an English template. ### Infinite Context Thanks to Phi-3.5 mini, NuExtract 1.5 has a 128k token context (~200 pages). But processing long documents with a transformer is memory-intensive: every token attends over every other token. Maxing out 128k tokens requires 1TB of GPU memory. Inference memory usage of NuExtract 1.5 as function of the number of tokens in the document. To solve this we adopted an original solution: we train NuExtract to extract information while being given previously extracted information . Example of continuation extraction. The output is obtained from the text, the template, and previously extracted information. This "continuation" ability allows us to process arbitrarily long documents by iteratively re-injecting the current state of information through a sliding context window — reminiscent of recurrent neural networks. Memory footprint is bounded by the window size (less than 30GB at 10k window size, irrespective of document size). Comparison of GPU memory requirements for using NuExtract with a full extraction window and with a 10k tokens extraction window. The downside is that the output is generated several times, and that performance can degrade if the sliding window is too small. ### Training & Results We train Phi-3.5 mini (3.8B) on our dataset to get NuExtract 1.5. We also train Qwen 2.5 0.5B to get a tiny English-only variant. #### English Performance On our 600-example, 12-problem English benchmark: - Zero-shot : NuExtract 1.5 is quite better than the original NuExtract, and slightly better than GPT-4o. - Many-shot (45 examples per problem, fine-tuning for NuExtract, in-context for GPT-4o): GPT-4o is slightly better, but not by much. The big gap between NuExtract 1.5 and the tiny variant suggests that a bigger NuExtract could largely beat GPT-4o . Zero-shot results on the structured extraction benchmark. NuExtract 1.5 is slightly better than GPT-4o. Many-shot results on the structured extraction benchmark. GPT-4o is a slightly better than NuExtract 1.5. A 500× smaller model rivaling a frontier LLM is surprising. Three reasons: focus on a single task lets NuExtract redirect weights to text understanding; training enforces strict template compliance and JSON-only output; training drastically reduces hallucinations by forcing extraction from the input and teaching empty results when needed. #### Multilingual Performance NuExtract 1.5 is much better than the original NuExtract, but GPT-4o is still better in this case. Model size matters a lot for multilinguality — we expect to fill this gap with a bigger NuExtract. Multilingual zero-shot results on the structured extraction benchmark. #### Long Documents On 8k-10k token documents (~20 pages), NuExtract 1.5 beats GPT-4o . On 10k-20k token documents with a 10k extraction window, NuExtract 1.5 still beats GPT-4o . We need to go down to a 2k window for NuExtract 1.5 to become worse than GPT-4o — and even then it remains much better than the tiny variant. Performance on long documents (between 8k and 10k tokens). NuExtract 1.5 beats GPT-4o! Performance on even longer documents (between 10k and 20k tokens). NuExtract 1.5 beats GPT-4o while only using a 10k extraction window. Performance of NuExtract on long documents as function of the size of the extraction window. The continuation procedure works. ### Let's Use It! You can try NuExtract 1.5 here . Don't hesitate to give us feedback to help us improve the next versions :) # NuExtract: A Foundation Model for Structured Extraction — NuExtract Blog https://about.nuextract.ai/blog/nuextract-a-foundation-model-for-structured-extraction Back to blog ## NuExtract: A Foundation Model for Structured Extraction Alexandre Constantin Machine Learning Scientist Liam Cripwell Machine Learning Scientist Etienne Bernard Co-Founder & CEO June 24, 2024 We introduce NuExtract, a lightweight text-to-JSON LLM. NuExtract allows to extract arbitrarily complex information from text and turns it into structured data. This model can be directly used in a zero-shot setting or fine-tuned to solve a specific extraction problem. As usual, we open-source it under MIT license for everyone to use. ### TLDR We trained language models from 0.5B to 7B parameters on an LLM-generated structured-extraction dataset. The resulting models — NuExtract-tiny , NuExtract , and NuExtract-large — achieve similar or higher extraction performance than popular LLMs that are 100 times larger. Comparison of NuExtract models with popular generic LLMs in the zero-shot setting. NuExtract-large is at GPT-4o levels while being at least 100 times smaller. ### Structured Extraction Structured Extraction is the most general and versatile information extraction task: extract all kinds of information from a document — entities, quantities, dates, and so on — and identify their (potentially hierarchical) relationships. The extracted information is structured as a tree following a template (a.k.a. schema ), almost always expressed in JSON. Structured extraction toy example. There are 5 kinds of entities and two kinds of relations organized in a tree of depth 4. Even simple examples (5 entity kinds, 2 relation types, depth 4) are a headache for traditional information extraction methods. Real applications involve multi-page documents and deeper trees, challenging even for modern LLMs. We encountered two main families of applications: parsing technical documents (medical, legal, financial — increasingly to power RAG knowledge bases), and chatbot conversations (extracting the right info to make API calls in real time). In a sense, structured extraction is the holy grail of information extraction . ### Why Not Just Use GPT-4? GPT-4 with a good prompt can do structured extraction, but: - Performance saturates with in-context learning (as we showed in our NuNER paper ). - It is massive, expensive, and requires sharing data . Parsing of a chemical reaction from its description. Data provided by Iktos.ai. GPT-4 result: chemical substances and quantities extracted, but both reagents are misclassified. Average performance of various NER models as function of training size. In-context learning quickly saturates. To solve all this, we need a compact task-specific foundation model . ### Task-Specific Foundation Models A task-specific foundation model is specialized for a generic task — sentiment analysis, entity recognition, structured extraction — but agnostic in terms of data domain and specific problem. They are small, usable in private settings, and often better than much larger generic models at the task. The recipe: - Take a diverse corpus (e.g. C4). - Annotate it using a modern LLM with a proper prompt — annotations don't have to be perfect. - Fine-tune a compact generic foundation model on this synthetic data. NuExtract creation procedure. A generic small language model (Phi-3) is fine-tuned on synthetic data generated by an LLM (Llama 3) to obtain a model specialized in structured extraction. The result can be used zero-shot, with examples, or fine-tuned for a specific problem. ### Template Representation We represent the schema with a sort of empty JSON : { "reactants" : [{"name" : "" , "quantity" : ""}], "time" : [""] } Each array contains an element template; empty strings indicate fields to extract. Only string values — numbers can always be returned as strings. The format is simple, easy to read, and we believe examples are more informative than descriptions . ### Dataset Creation We use 300k English texts from the C4 dataset as base — diverse enough that we'll find something interesting to extract in most documents. We then prompt an LLM (Llama 3 70B) to generate a template from each text , with hand-crafted few-shot examples in the prompt. Once we have templates, we use the LLM again to extract information according to each template . For half the examples, we extract from the full text; for the other half, we remove parts of the text to teach the model that it's acceptable to return empty strings — a form of negative sampling that fights hallucinations . The extraction prompt forces the LLM to copy-paste from the text rather than generate new values — another anti-hallucination tradeoff. After filtering for template compliance and value-in-text checks, we end up with 50k annotated examples . Typical example from C4 annotated by Llama 3 70B. 16 words, extraction depth of 5. Word counts mostly stay below 200, with a tail up to 1200 (2-3 page documents). Extraction-tree depths span 3-5 (vs. depth 1 for classification/NER, depth 2 for relation extraction). The LLM produced 200k+ unique field names , with healthy domain coverage (dates, contact info, dimensions, nutrition, health, …). Distribution of the number of words across all text pieces. Distribution of extraction-tree depths across all examples. Word cloud of the 100 most common field names found by Llama 3 70B. Feature map of the 10k most common field names. We also add a hybrid few-shot setting : 0-3 output examples (no input) added to the prompt. NuExtract can be used pure zero-shot or with output-only examples. ### Base Models Structured extraction has a large complex output space, so we need to generate the output. We use pure decoder LLMs: - Phi-3-mini (3.8B) → NuExtract - Phi-3-small (7B) → NuExtract-large - Qwen1.5-0.5B (0.5B) → NuExtract-tiny ### Evaluation NuExtract result on a toy example. All values are correct except for one that needed generation. We built a dedicated benchmark from a set of "problems" (e.g. resume parsing) with hand-extracted ground truths. We use a tree-matching metric that aligns extracted leaves recursively, scoring exact value matches between 0 (different) and 1 (perfect). Benchmark to be released publicly when finalized. Zero-shot results : - NuExtract-tiny is better than GPT-3.5 while at least 100× smaller. - NuExtract outperforms Llama3-70B while 35× smaller. - NuExtract-large reaches GPT-4o levels while at least 100× smaller. Small specialized models 100× smaller than frontier LLMs bring three benefits: lower inference cost , local/private deployment , and easy fine-tuning . After fine-tuning (5-fold cross-validation on a 50-example chemistry problem from Iktos.ai): - NuExtract-tiny (0.5B) becomes a bit better than GPT-4o. - NuExtract and NuExtract-large reach a different level entirely. These results show the benefits of using small, fine-tuned, task-specific models for structured extraction. ### Let's Use It! Structured extraction is one of the main use cases of modern LLMs. NuExtract reaches similar or higher performance than the largest LLMs while being orders of magnitude cheaper to use. MIT-licensed, available for everyone. For even higher performance, talk to us 🙂. # A Foundation Model for Entity Recognition — NuExtract Blog https://about.nuextract.ai/blog/a-foundation-model-for-entity-recognition Back to blog ## A Foundation Model for Entity Recognition Sergei Bogdanov Machine Learning Scientist Alexandre Constantin Machine Learning Scientist Etienne Bernard Co-Founder & CEO November 7, 2023 Entity recognition is a widely used information extraction task, yet publicly available foundation models are not well suited for it. We leverage modern LLMs to create a small-yet-powerful foundation model for this task. This BERT-size model can be used to create custom entity recognizers with typically 5x less annotated data than before. This model is powering NuMind and we open-source it with an MIT license for everyone to use. ### Entity Recognition Entity recognition (a.k.a. NER) is the task of detecting entity types and other concepts mentioned in text — cities, people, companies, or any instance of human concepts. It powers automatic news analysis, medical coding, legal document analysis, and many other applications. Legal document annotated with entities (from NuMind's annotation interface). You could try GPT-4 with a good prompt — fine for some cases. But for high volumes, better performance, or confidentiality , the traditional deep learning approach (fine-tuning a BERT-like foundation model on hand-annotated data) is preferable. The catch: you typically need hundreds of annotated documents to reach (and surpass) GPT-4 performance. Two ways to reduce that effort: annotate automatically with an LLM, or use a better BERT-size foundation model . This post is about the second approach. ### RoBERTa Out-of-the-Box Transformers like BERT/RoBERTa are pre-trained self-supervised on missing-word prediction. The last-layer embeddings carry contextual semantic information — visible when "Amazon-the-river" and "Amazon-the-company" cluster differently. Visualization of BERT embeddings. From: Introduction to Machine Learning. The classic approach is to attach a linear classifier on top of these embeddings to predict, per token, whether it belongs to a given concept. Neural network computing the token probabilities to be part of a particular concept. We evaluate on MIT Movie , MIT Restaurant , OntoNotes 5 , and BioNLP 2004 , training only the linear classifier (other layers frozen). This is our baseline. Transfer learning performance using the last layer of English RoBERTa base. ### Leveraging Human-Annotated Datasets Last-layer embeddings carry contextual word meaning, but there's no reason for them to encode human concepts in a linearly accessible way . We want a foundation model that explicitly knows about human concepts , with this information surfaced in the last layer. The simple idea: train an entity recognizer on a dataset annotated with a large, diverse set of concepts. Some researchers did this using the NER Corpus (16M examples, 475M tokens, 315 unique concepts from Wikipedia). Fine-tuning the last 6 layers of RoBERTa on it gives a small improvement over the baseline — better in the few-shot regime, equivalent for larger training sets. Transfer learning performance of RoBERTa base, and RoBERTa base fine-tuned on NER Corpus. A 315-concept set is small. Wikipedia is one domain. We want tens of thousands of concepts on highly diverse data . Human annotation would be too costly. Modern LLMs solve this. ### Using LLMs to Annotate Human Concepts Annotating from a pre-defined ontology gives mediocre results, even with GPT-4. Our key insight: let the LLM figure out the ontology as it annotates , introducing concepts on the fly. This gives both ontology diversity and high annotation quality — even GPT-3.5 works in this setting. The prompt asks the model to "label as many entities, concepts, and ideas as possible", invent new entity types where needed, and output entity from the text -|- entity concept -|- description . We applied this to 160k English sentences from the C4 dataset , producing about 800k annotations and 80k unique concepts — long-tailed distribution, with the 100 most common concepts accounting for 43% of annotations and ~50k rare concepts appearing once. Concept counts, sorted from most to least common. Cumulative concept counts. The 100 most common concepts include classics ("person", "location") and richer ones ("medical procedure", "operating system"). 28k unique words appear in concept names. This dataset has much higher concept diversity than any available human-labeled dataset . 100 most common concepts present in the dataset, size reflects their frequency of appearance. Sample of concepts appearing only once in the dataset. Feature plot of the 80k unique concepts present in the dataset. Annotations contain mistakes but few false positives — the main issue is false negatives (missed concepts). As we'll see, this isn't a problem for the foundation model. ### Learning From LLM-Annotated Concepts Naive approach (linear classifier on top of RoBERTa) doesn't work — too many concepts, many similar, most rare. Our trick : instead of independent weight vectors per concept, compute the weight vector as f(concept_name + description) where f is a sentence encoder. This lets the network leverage similarities between concepts and scales to many more concepts. Neural network computing the probability for each token to be part of a particular concept. This network is trained on our dataset to create the foundation model. In practice, we ignore concepts not present in the current batch — we don't learn from "negatives" of other concepts. Probabilities become uncalibrated, but we only care about token embeddings. False negatives in the data also become a non-issue. The setup is essentially contrastive learning . We use RoBERTa for both networks and fine-tune the last six layers. After training, we keep only the token-encoder network — that's our foundation model. ### Results Compared to RoBERTa-base and RoBERTa-fine-tuned-on-NER-Corpus, our foundation model is much better both in few-shot and as data scales . F1 differences are large, but the better metric is data efficiency : - ~30 examples per concept needed for previous models to reach F1=0.65 - ~5 examples per concept needed for ours — 6× more data efficient in this regime Transfer learning performance of RoBERTa base, RoBERTa fine-tuned on NER Corpus, and RoBERTa fine-tuned on our dataset. Data efficiency keeps increasing with training size. Per-dataset breakdown shows our model is substantially superior across all datasets and data regimes . Best case: >10× data efficiency on BioNLP2004 / MIT Movie. Even on the favorable MIT Restaurant: 2-3× improvement. Per-dataset transfer learning performance. The large gain (compared to the moderate gain from training on NER Corpus) likely comes from the combination of a large diverse concept set, domain-diverse data, and the specific training procedure . ### Let's Put It to Work This foundation model is an important step forward, allowing accurate entity recognizers to be trained with substantially less annotated data. We open-source it under MIT license: - English model - Multilingual model Of course, the best way to use it is through NuMind 🙂. # Creating Task-Specific Foundation Models with GPT-4 — NuExtract Blog https://about.nuextract.ai/blog/creating-task-specific-foundation-models-with-gpt-4 Back to blog ## Creating Task-Specific Foundation Models with GPT-4 Alexandre Constantin Machine Learning Scientist Sergei Bogdanov Machine Learning Scientist Etienne Bernard Co-Founder & CEO August 25, 2023 There are two kinds of BERT-size NLP models in this world: general-purpose ones (a.k.a. foundation models), and highly specialized ones, trained on specific tasks and data. Neither kind is ideal to solve particular NLP problems on your data. We need to fill this specialization gap with task-specific foundation models, and propose a way to create them efficiently using LLMs. We apply this method to create a state-of-the-art domain-agnostic foundation model for Sentiment Analysis that we open source for everyone to use. ### Introduction → Check all our task-specific foundation models and datasets here . Let's say that you want to know the sentiment of your users by analyzing their messages on your platform. You might want to know if they are going to churn or something like that. You could go the LLM route, but it is expensive. The alternative is to use a good old BERT-size model: they are small, fast, and powerful enough to solve most text classification tasks**.** To proceed, you would start from a pre-trained foundation model that is trained in an unsupervised way on a diverse dataset, and then fine-tune it on your task and data — a classic transfer learning procedure. However, fine-tuning requires to label many examples in order to reach the desired level of specialization. Alternatively, you could take a model that is already trained on a sentiment analysis task , such as something trained on tweets sentiments , and fine-tune this model instead. Unfortunately, your data is qualitatively different from tweets, and what you mean by "sentiment analysis" also differs a bit. So, in truth, these specialized models are not so helpful , unless you are lucky enough for your task and data to be very similar to the pre-trained model's task and data. In this case, a much better solution would be to start from a model that is pre-trained on a broad definition of "sentiment analysis", and on a variety of domains (tweets, blogs, news, reviews, emails, etc.). We call these models task-specific foundation models to highlight that they are broadly specialized to a kind of task, but domain-agnostic, and intended to be used as foundation models. Note that there already are public domain-specific foundation models and language-specific foundation models, but task-specific foundation models are lacking. Fortunately, we came up with a simple, cheap, and reliable way to create these models using LLMs , and thus efficiently filling up this specialization gap. In the following we show how to create these task-specific foundation models from diverse data that was automatically labelled by an LLM , using sentiment analysis as an example. The resulting foundation model achieves much better transfer-learning performance than other pre-trained models, achieving in some cases >10x data efficiency, such as in the case of this financial news dataset: Transfer learning performance on the financial_phrasebank dataset for our sentiment analysis foundation model compared to a state-of-the-art generic foundation model (e5-base-v2). ### Task-Specific Foundation Model Ok, we want to train a foundation model that is good at a particular task — in this case, sentiment analysis — and easily adaptable to all kind of domains. The simplest approach is to take existing labelled datasets for this task and train a model on a mix of these datasets. The resulting model can then be fine-tuned on a new dataset. The issue with this method is that we are limited by the diversity of public datasets . Indeed, almost all public datasets for sentiment analysis either consist of tweets, product reviews, movie reviews, or financial news. While these domains are definitely common for sentiment analysis, there is a long tail of other domains that need of sentiment analysis, and we have seen a few of them with our customers (group chat messages, internal user messages, etc.). Also, public datasets are often created for academic purpose and with preprocessing such as language filtering that harms the diversity of the data. To create a truly domain-agnostic, task-specific foundation model, we would like to train on all kind of data , and even with slightly different meanings of "sentiment analysis". Unfortunately, there aren't diverse-enough labeled datasets for this purpose, but we can create one! Recent LLMs like GPT-4 or even GPT-3.5 (especially the earliest version) are very good at understanding text and can be used to label data automatically . This means that we only need to find a large and diverse dataset, write a good prompt for an LLM to automatically label the data, and then train a model on this data to obtain our task-specific foundation model . To do so, we randomly select 300,000 text snippets from the C4 dataset , a large general-domain dataset from the web, which we annotate using GPT-3.5 with the following simple prompt: "The goal is to create a dataset for sentiment analysis. Classify the input text as Positive, Negative, or Neutral. Return only the label. Do not return the input text or anything else." Note that we don't explicitly mention what is "Positive", "Negative", and "Neutral" on purpose, to allow for different interpretations by the LLM. This ensures that the sentiment analysis foundation model is not overspecialized to a specific kind of sentiment analysis. We tried more complex prompts, but this simple one obtained great results. Here is a sample from the annotated dataset: Text Label Another 400 charities are in danger of losing their status with the national charity regulator for failing to make contact with them. Negative Happiness is sweet persimmons wrapped with brie and ham in a buttery puffed pastry Positive From your physical assets down to business data that are critical to the operation of your business – all of it can be found within your commercial area or space Neutral Unsurprisingly, this dataset is class-imbalanced. Most documents are neutral or positive, with only 5% of them being negative. Class balance of the LLM-annotated sentiment analysis dataset. Learning from such imbalanced data requires unnecessary computation, so we randomly remove documents from the "Neutral" and "Positive" classes to obtain a balanced dataset. In the end, we obtain 40,000 annotated sentences that you can download here . Now, we just need to train our task-specific foundation model on this dataset. To do so, we start from the e5-base-v2 model , which is a state-of-the-art BERT-size foundation model. Then, we fine-tune its last three layers , as it proved to lead to higher performance, better model stability, and higher data-efficiency. We obtain a pretty good generic sentiment analysis model that you can download here . ### Performance Evaluation We are not interested in directly using our model to analyze sentiment, but rather using it as a foundation to be adapted to specific tasks and domains. Therefore, in order to test the transfer learning abilities of our foundation model, we use four sentiment analysis datasets from typical domains : airlinetweetSA , financial_phrasebank , amazon_en , and climate_en , and compare the performance against three alternative foundation models : e5-base-v2 (E5), e5-base-v2 fine-tuned on the SST2 general-domain sentiment analysis dataset (E5-SST2), and e5-base-v2 fine-tuned on all evaluation datasets except the one currently being evaluated (E5-Multi). All of these models were obtained by fine-tuning last three layers. Note that, while E5 is a generic foundation model, E5-SST2 and E5-Multi are already task-specific foundation models . To evaluate the performance of each foundation model, we fine-tune it on some examples from a dataset and measure the resulting model's performance on the corresponding test set . To simplify the fine-tuning process — and because we are only interested in comparing the performance of foundation models — we simply train a logistic regression on top of the last layer of the foundation model instead of a full fine-tuning procedure. Here is the performance we obtain for the F1-score: Dataset # Examples E5 E5-SST2 E5-Multi Ours airlinetweetSA 1 0.360 0.488 0.454 0.540 * airlinetweetSA 5 0.491 0.574 0.628 0.651 * airlinetweetSA 10 0.546 0.616 0.648 0.680 * airlinetweetSA 50 0.675 0.685 0.703 0.715 * amazon_en 1 0.234 0.314 0.355 * 0.327 amazon_en 5 0.352 0.407 0.441 * 0.426 amazon_en 10 0.382 0.430 0.448 0.455 * amazon_en 50 0.437 0.461 0.484 0.486 * climate_en 1 0.444 0.443 0.453 0.516 * climate_en 5 0.622 0.634 0.643 0.667 * climate_en 10 0.685 0.681 0.684 0.692 * climate_en 50 0.741 0.741 0.750 0.751 * financial_phrasebank 1 0.321 0.427 0.478 0.585 * financial_phrasebank 5 0.485 0.576 0.643 0.692 * financial_phrasebank 10 0.540 0.624 0.671 0.732 * financial_phrasebank 50 0.664 0.695 0.720 0.749 * And here is a visualization of these results: Transfer learning performance of foundation models on various sentiment analysis datasets. There are two important takeaways from these results. First, all three task-specialized foundation models (E5-SST2, E5-Multi, and our LLM-annotated-dataset model) beat the generic E5 model . This shows the utility of using task-specialized foundation models, and is not that surprising. Second, something more surprising, our LLM-annotated-dataset foundation model is clearly better than the other task-specialized foundation models . This shows the superiority of our approach to creating such a model compared to using existing dataset. The LLM used might not annotate as well as a human, but this is largely compensated for by the diversity of data found in C4 and the diversity of the "sentiment analysis" definition implied by the prompt. In the end, there is more generic knowledge about the sentiment analysis task inside our LLM-annotated dataset than inside available human-annotated datasets, even if they claim to be generic, and even if they are combined! ### Let's Fill The Gap! We found a simple method for creating high-quality task-specific models using modern LLMs. We applied this method to create a state-of-the-art, domain-agnostic foundation model for sentiment analysis. We have released this model with an open source license , and included it in NuMind so that all of our users can efficiently create custom sentiment analysis models. Now, given the success of this method, we think it is time to create additional task-specific foundation models for other tasks (such as toxicity detection or topic identification) and other languages. These models will allow to quickly obtain customized BERT-size models for all kind of domains , hence democratizing NLP further. Let's fill this specialization gap! # What are Large Language Models? — NuExtract Blog https://about.nuextract.ai/blog/what-are-large-language-models Back to blog ## What are Large Language Models? Etienne Bernard Co-Founder & CEO June 21, 2023 These past few months, thanks to ChatGPT and its siblings, we have been witnessing something historic. It seems that computers are finally able to understand our language, and are even able to speak back! These AIs are the latest iterations of large language models , also known as LLMs . But what exactly are these LLMs? How do they work? And how are they created? Let's dive into it. ### Language Models In a nutshell, a language model is something that is able to generate text in some way . Language models have plenty of applications. For example, you can use them to analyze sentiment, flag toxic content, answer questions, summarize documents, and so on. But in principle, they could go far beyond these usual tasks. Indeed, imagine, for example, that you have a perfect language model, something that can generate any kind of text in such a way that it is impossible to distinguish whether this text is generated by a computer or not. Then, you could do plenty of things with it. For example, you could make it generate classic content such as emails, news articles, books, and movie scripts. But then you could go a step further and make it generate computer programs or even entire software. And then, if you are really ambitious, you could make it generate scientific articles. If the language model is truly "perfect", these scientific articles would be indistinguishable from real articles, which means the language model would have to conduct actual research! Of course, such a perfect language model is out of reach at the moment, but this gives an idea of the potential power of these systems. Language models are not "just predicting text"; they are potentially much more than that. Let's now look at what these models are in practice, starting from the first kind of naive language models to the current transformer-based large language models. ### Naive Language Models Language models are machine learning models, which means that they learn how to generate text. The way to teach them (a.k.a. the training phase) is to give them a large corpus of text, from which they figure out how to imitate the generative process that created it . Ok, this is rather abstract, but it is actually easy to create a naive language model. You can take a corpus of text, chunk it into strings of a certain size, and measure their frequencies. Here is what I got with strings of size 2: From Introduction to Machine Learning These chunks are called *n-grams (*where n is their size, so n =2 here). From these n -grams you can generate text by playing dominoes. You start with an initial n -gram, let's say "th", and then randomly select - according to the measured frequencies - one n -gram whose beginning matches the end of the initial n -gram. Here it could be "hi", which would make "th"+"hi"= "thi". You can then continue by attaching an n -gram starting with a "i", and so on to generate entire text. As you probably guessed, these n -gram models do not generate the most coherent text. Here is what I got when continuing the procedure: thint w dicofat je r aton onecl omitt amen h s askeryz8, orbexademone ttexind thof thevevifoged tc hen f maiqumexin sl be mo taicacad theanw.soly. fanitoila, al Not great, to say the least! This makes sense because the model only takes into account the previous character to make its next-character prediction - it has a tiny memory. If we use n=4, we get something slightly better: complaine building thing Lakers inter blous of try sure camp Fican chips always and to New Semested and the to have being severy undiscussion to can you better is early shoot on Now there are some correctly spelled words, but this is still not great! In theory, increasing n further will make things better, but in practice, we cannot increase n much without requiring a gigantic dataset to train the model on . One last thing we could do is to use words instead of characters as the base unit (the base unit is called token in NLP jargon). It will improve things, but it won't lead to very coherent text either since we are limited to n <6 . These naive language models always have a short memory and thus cannot generate coherent text beyond a few words . They do have some use cases, though. Until a few years ago, they were used extensively for text classification and speech recognition, and they are still used today to identify languages, for example. However, for more advanced text understanding and text generation tasks, these models are not sufficient. We need neural networks! Modern language models are based on (artificial) neural networks. Neural networks are brain-inspired computing machines that are able to learn how to perform a task from examples of that task . This form of machine learning is also called deep learning because the networks are composed of several computational layers (hence they are "deep"). In a neural network, learning is done by going through the examples of the task and iteratively modifying the parameters of the network to optimize the task objective . You can think of these parameters as a bunch of knobs that you can turn left and right to improve the objective , except that it is the computer turning them for you, and it knows how to turn them all at once in the correct directions to improve things (thanks to the famous backpropagation algorithm ). So the network goes through the examples of the task (typically by batch of a few hundred examples) and optimizes the objective as it goes. Here is an example of an objective (called a cost function, the smaller the better) being optimized: ### Neural-Network Based Language Models Cost function as function of training iterations. From Introduction to Machine Learning. As the model is trained, the cost goes down, which means that the model gets better at its task. Ok, so in our case, we want to generate text. The current standard way to do this is to train a model on the task of predicting the next word from previous words . Since there are several possible continuing words, the model learns to associate a probability with every possible continuing word. Here is a visualization of this probability distribution for what comes after "the cat sat on the": Modified from Introduction to Machine Learning Once we have such a predictive model, we can generate a word by sampling from its predicted probabilities, then re-inject this word into the network, sample a new word, and so on to generate text. Such a generative procedure is said to be autoregressive because the network reuses its own predictions to make new predictions. There are arguments as to why this is not the best way to do it, but it is currently how it is done, and it works surprisingly well! Historically, the first kind of neural networks that have been used to model language is recurrent neural networks (a.k.a. RNNs ), and in particular the LSTM architecture. RNNs are like little machines that read words one by one, updating a sort of mental state in the process. This mental state is like a memory, and allows them to - potentially - remember things for a while as they generate text. RNNs can obtain much better performance than n -gram models. The RNN that impressed me the most at the time of its creation was A Neural Conversational Model , which features an LSTM trained by Google in 2015 on movie subtitles. This model is, by all means, already a large language model (maybe be the first?): it has more than 300 millions parameters, and has been trained on about 1 billion words. Since movie subtitles are conversations, you can readily use this model as a chatbot. Here are examples of answers from various human inputs: Human: what 's your name ? Machine: i 'm julia . Human: when were you born ? Machine: july 20th . Human: my name is david . what is my name ? Machine: david . Human: who is skywalker ? Machine: he is a hero . Human: what is the color of the sky ? Machine: blue . Human: what is the color of blood ? Machine: it is the same as a black eye . As you can see, the model is able to chit-chat, but it also knows a little bit about the world, which it learned solely from learning to predict text! I remember being fascinated by this fact: learning to predict text forces you to understand the world (which does not mean it is easy by any means). However, this model has strong limitations. It is often wrong and, like similar LSTM-based models, cannot generate long coherent texts. Indeed, in theory, RNNs can remember things for a long time, but in practice, they tend to forget things fairly quickly: past a few dozen to a hundred words, they start to derail and become incoherent . One solution to this short-term memory issue came in 2017 from a new kind of neural network called transformers , which is based on the attention operation (which is essentially a selection operation). As an eye candy, here is how the transformers are depicted in their introductory paper for the task of translation: Transformer architecture. From https://arxiv.org/abs/1706.03762 There are plenty of interesting things to say about this architecture, but the bottom line is that transformers works very well for modeling text, and it is well adapted to be run by graphics cards (GPUs) in order to process (and learn from) large amounts of data. It is this transformer architecture that led to (or at least strongly contributed to) the emergence of modern large language models . ### Modern Large Language Models The invention of transformers marked the beginning of the era of modern large language models. Since 2018, AI labs have started to train increasingly larger models. To the surprise of many, the quality of these models kept improving! Here is a visualization of these models, from which we will highlight the notable ones: Evolutionary tree of LLMs. From https://github.com/Mooler0410/LLMsPracticalGuide There are three main flavors for these language models. One type (shown in pink on the picture, the "encoder-only" group) includes LLMs that are good at text understanding because they allow information to flow in both directions of the text. Another type (shown in blue in the picture, the "decoder-only" group) includes LLMs that are good at text generation because information only flows from left to right of the text in order to generate new words efficiently in an autoregressive fashion. Then there is an encoder-decoder type (shown in green) which combines both aspects and is used for tasks that require understanding an input and generating an output, such as translation. It mostly started with the text understanding kind. First with ELMo (still using RNNs) and then the famous BERT from Google, and its descendants like RoBERTa , which are all transformers . These models typically have around a few hundred million parameters (corresponding to around 1GB of computer memory), are trained on around 10GB to 100GB of text (so typically a few billion words), and can process a paragraph of text in about 0.1s on a modern laptop. These models have drastically improved the performance of text-understanding tasks such as text classification, entity detection, and question answering . This was already a revolution in the field of NLP, but it was just the beginning... In parallel with the development of text-understanding LLMs, OpenAI began creating text-generating LLMs based on transformers . First, there was GPT-1 in 2018 , which had 100 million parameters, **and then GPT-2 in 2019 , which has up to 1.5 billion parameters and is trained on 40GB of text. The creation of GPT-2 was, at least to me, a pivotal moment.** Here is the kind of text it can generate, starting from a human-written paragraph: GPT-2 demo. From https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf This is excellent English, and the text is coherent. For example, the name of the scientist does not change, which would be a classic issue with RNN-based models. GPT-2 was such a leap in generation quality that OpenAI originally decided not to release it to the public for fear of harmful use . GPT-2 was a sign that LLMs were on the right track. Note that the way to use such a language model is to give it a starting text to be completed. This initial text is called a prompt . One year later (2020), **OpenAI created GPT-3 , a model with 175 billion parameters** (700GB of computer memory to store the model!). This was a significant increase in size, and it represented another significant improvement in terms of text generation quality . In addition to its improved performance, GPT-3 has been eye-opening in terms of how we might use LLMs in the future. First, GPT-3 is capable of writing code . For example, you can use it to generate (very) simple websites by describing what the website should look like in the prompt. Here is an example where we ask GPT-3 to create a button in HTML: Screenshot of a GPT-3 prompt asking for an HTML button, next to the HTML code the model generated in response These basic coding abilities were not so useful at the time, but they hinted that software development could be radically transformed in the future . Another eye-opening insight from GPT-3 is that it can perform in-context learning , which means it has the ability to learn how to perform a task by only being shown examples in a prompt. This means that you can customize these LLMs without having to change their weights, just by writing a good prompt. This has opened up a new kind of NLP, purely based on prompting, which is now very popular. Overall, GPT-3 revealed the potential of prompting as a new way to make machines do what we want them to do through natural language. Note that GPT-3 is much larger than GPT-2. Since 2018, we have witnessed an extreme increase in model sizes. Here are some notable LLMs, along with their sizes: Chart comparing the parameter counts of notable large language models, rising from GPT-1 to models approaching one trillion parameters In two years, the number of parameters has been multiplied by 1000, and the current largest models ( like GPT-4 ) are close to 1 trillion parameters. This increase was driven by the fact that performance kept on improving with model size, with no plateau in sight. These models are so big that we might be tempted to compare them with our brain, which has around 100 billion neurons, each connected to around 1,000 other neurons on average, so about 100 trillion connections in total. In a sense, the largest LLMs are still 100 times smaller than our brain. Of course, this is a very loose comparison since our brain and current LLMs use very different architectures and learning procedures. Another interesting metric about these models is the number of words that they "read" during their training phase: NB: This shows the number of words processed by the network and not the number of words in the dataset (since a dataset can be only partially used or used several times during the training procedure). As you can see, it is a lot. These models see more than 100 billion words during their training, which is more than 100 times what a human will ever hear or read in their lifetime! This shows how different these neural networks are from our brain. They learn much more slowly than us, but have access to much (much!) more data. Note that the number of words that LLMs encounter during their training did not increase as much as the parameter count (only a factor of 3 between GPT-1 and GPT-3). This is because model size was prioritized instead, and it turned out to be a bit of a mistake. The latest models are not much larger than GPT-3, but they are trained by processing much more words than GPT-3. The issue with this hunger for data is that there is a hard limit on the total amount of useful text available - a few trillion words - and models are getting close to it. There is still the possibility to loop over all this text, but this results in diminishing returns in terms of model performance . Overall, we can consider that there is an effective limit of a few tens of trillions of words to be processed by the network during its training phase - about 10 times more than GPT-4 experienced. The other issue, which arises from training larger models on more data, is that the cost of computing is increasing. Here are the estimated computation costs for training the models mentioned above: Chart of estimated training compute costs for major large language models, increasing sharply with model size To significantly outperform current models, the next generation of models should require hundreds of millions of dollars in computation , which still makes sense given the benefits these models provide, but is an issue nonetheless. Scaling models up is becoming increasingly difficult. Fortunately, scaling up is not the only way to improve LLMs. At the end of 2022, an innovation unlocked yet another revolution, with an impact far beyond the world of NLP this time. ### Instruction-Tuned & Chatbot LLMs GPT-3 revealed the potential of prompting, but writing prompts is difficult. Indeed, classic LLMs are trained to imitate what they see on the web, so to create a good prompt you have to figure out what would be, on the web, the initial text that would lead to your desired output. This is a weird game and kind of an art to find the right formulation. You need to change the wording, pretend that you are an expert, show examples of how to think step by step, and so on. This called prompt engineering , and it makes using these LLMs difficult. To address this, researchers have been exploring how to modify these base LLMs to better follow human instructions. There are two main ways to do this. The first is to use instruction-answer pairs that are written by humans and then fine-tune (i.e., continue training) the base LLM on this dataset. The second way is to have the LLM generate several possible answers, have humans rate these answers, and then fine-tune the LLM on this dataset using reinforcement learning. This is known as the famous Reinforcement Learning from Human Feedback (RLHF) procedure. It is also possible to combine both approaches, which is what OpenAI did with InstructGPT and then with ChatGPT. Instruction-tuning step of InstructGPT and ChatGPT. From https://openai.com/index/chatgpt/ (modified from https://arxiv.org/abs/2203.02155) Using both techniques together results in an instruction-tuned LLM that is much better at following human instructions than the base model, and therefore much easier to use. Instruction-tuned LLMs were already great, but there was one last step to turn these LLMs into something that could truly be used by everyone: making a chatbot version of them. OpenAI achieved this by releasing ChatGPT in December 2022, a chatbot based on GPT-3.5. It has been created in the same way as InstructGPT, but this time using entire conversations instead of just instruction-answer pairs. After the release of ChatGPT, we witnessed a number of new LLM-based chatbots. OpenAI improved ChatGPT by using GPT-4 instead of GPT-3.5, Anthropic released Claude , Google released Bard , Meta released LLaMA , and several open-source LLMs are currently being released. This is a real explosion, and I believe it will lead to many exciting applications - something that we, at NuMind, will help with. Two months after its release, ChatGPT already had 100 million users, the fastest product growth ever. People use it to write emails from bullet points, to reformulate text, to summarize text, to write code, or just to learn something - a task that search engines had the monopoly of until then. The release of ChatGPT was a turning point in the history of LLMs. Everyone realized the potential of these LLMs, and an "AI race" started, involving the main AI labs in the world and several startups. Note that the sudden widespread accessibility of LLMs also comes with the concern that they will be used to do harmful things. This is why a big part of creating these open-ended LLM-based chatbots is about making them "safe" (or "aligning them with human values"), which means that they should not help you build a bomb, for example. At the moment, there are often ways to trick the chatbots and bypass their safeguards, but these safeguards are getting better over time. I believe that it will become very hard to trick them eventually. ### What's Next? LLMs have improved a lot these last years, and there is more effort than ever directed at improving them further. So, what should we expect for the next few years? It is hard to predict the future, but here are some thoughts. One obvious direction is to continue scaling up model sizes and the amount of training data. This has worked extremely well in the past and should still allow for some improvements. The issue is that training costs are becoming prohibitive (>$100M). Better GPUs and new specialized hardware will help, but they take time to be developed and produced. Also, the biggest models already iterate over all books and about the entire web, which means we are reaching the limits in terms of available training data (the so-called "token crisis"). So, for sure, there will not be an explosion of parameter numbers in the next few years like we saw in the last few years. The largest models should settle below 1 trillion parameters this year, and then experience something like a 50% annual growth at most. Another obvious direction is to go beyond pure language models and incorporate images or even videos into the training data - that is, to train multimodal models. Learning from such data might help these models understand the world better. GPT-4 has been trained on images as well as text, and it improved performance a bit (but not so much). Training on videos might change the game, but it requires a lot of computation. I would expect us to have to wait 2+ years before seeing the first real large "language" model trained on videos. Scaling up or going multimodal will require a lot of computation. A solution to mitigate this issue is to use better neural architectures and training procedures that are either less computationally intensive or that can learn with less data (and our brain is proof that it is possible). Most likely, RNN-like memory will make a comeback because it is so efficient at runtime (see for example the recent RWKV architecture ). But we could also see a more drastic change, such as LLMs that do not generate in an auto-regressive fashion but in a top-down fashion - basically making (random) decisions prior to generating words - which seems like a more logical thing to do when you think about it (and is how neural networks generate images at the moment). It is hard to know when such new architectures/methods will be developed, but I would not be surprised if it happens in the next few years and leads to greatly improved LLMs. One other direction for improvement is to follow up on the instruction-tuning route and involve many more humans in "educating" the LLM (a.k.a. aligning the AI). This could be done by private AI labs, but it could also be a more crowd-sourced Wikipedia-like project to improve and align LLM capabilities of open models. On that topic, we might also want to deviate from the traditional RLHF and have people just discuss with the model to teach it, as we would do with children. I'm not sure about the timeline for such a project, but I have been thinking about this for a while and would love to see it happen! Ok, we only talked about improving the actual model, but there are ways to improve LLMs without even changing the model. One such way is to give LLMs access to tools. Such a tool can be a search engine to find accurate information, or a calculator to do basic math. It can also be a knowledge base coupled with an inference engine (a classic component of symbolic AI) such as Wolfram Alpha to find facts and perform logical reasoning or other kinds of computations that neural networks are not great at. And of course, this tool can be a full-on programming environment to write and run code. LLMs can use these tools by generating special tokens (words) which trigger API calls and then inserting the API output in the generated text: Examples of an LLM using tools. From https://arxiv.org/abs/2302.04761 This tooling trend has already started (see e.g. ChatGPT plugins , the LangChain library, and the Toolformer paper ) and I believe it will become central to LLMs. Another direction is to use the LLMs in a smarter way so that they become better at completing tasks. This can be achieved through clever prompting or a more advanced procedure. One simple example of this is to ask the LLM to think step by step. This is called chain-of-thoughts prompting and improves the performance of LLMs on tasks that require logic . Here is an example of how to prompt an LLM to think step by step: Chain-of-thought prompting example. From https://arxiv.org/abs/2201.11903 Similarly, you can ask the LLM to reflect on its output, criticize it, and modify it in an iterative fashion. These kinds of iterative procedures can improve performance significantly, especially for generating code. Then, you can go even further and create fully autonomous agents that can manage a list of tasks and iterate over these tasks until the main goal is reached (see AutoGPT and BabyAGI ). These autonomous agents are not working well at the moment, but they will improve, and it is difficult to overstate how impactful they may become. By the way, since an LLM can improve its answers through these procedures (chain-of-thoughts, iterative critiques, etc.), we can create instruction-answer pairs using these procedures and then fine-tune the LLM on these pairs in order to improve its performance. This kind of self-improvement is possible (see, for example, here ) and I believe it has a lot of potential. We could, for example, imagine the model discussing with itself in order to become more self-consistent, a sort of self-reflection procedure. This direction will probably give another boost to LLM performance. Ok, I probably missed other directions for improvements, but let's stop here. Overall, we can't know for sure what the future holds, but it is clear that LLMs are here to stay. Their ability to understand and generate text makes them a fundamental piece of technology. Even in their current form, LLMs will unlock plenty of applications - the most obvious one being digital assistants that actually work - and in the craziest scenario, they might even lead us to the creation of some kind of super-intelligence - which is a topic for another time! # Seed Round Completed — NuExtract Blog https://about.nuextract.ai/blog/seed-round-completed Back to blog ## Seed Round Completed Etienne Bernard Co-Founder & CEO March 30, 2023 We are delighted to announce that we closed a seed funding round! We raised $3M, which we will use to continue building our NLP tool focused on text understanding. If your company deals with text of any kind (customer feedbacks, internal documents, code, etc.) chances are that you would gain to understand and process it automatically. We use Large Language Models (LLMs) - and a new learning paradigm that we call interactive AI development - so that data scientists, data analysts, and software developers can tackle these tasks efficiently. Ok, before diving further into what we do, we would like to give a big thanks to our investors. Flybridge for leading this round, Big Bets for putting the first institutional check, Carya, Pioneer, Velocity, Sharpstone, all our business angels for believing in us, and of course YCombinator for kickstarting everything. We look forward to working with all of you and make NuMind a leading AI company! Logos of the investors backing NuMind's seed round, including Flybridge, Carya, Pioneer, Velocity, Sharpstone and Y Combinator And now a bit more about NuMind… ### The Large Language Model Revolution Unless you have been living under a rock these last months, you heard about the rise of LLMs and their most famous representatives: ChatGPT and GPT-4. These models - trained on a big chunk of the web - are able to generate text at a level that was unthinkable a few years ago. This opens up plenty of new applications, notably in text generation, natural language interfaces, but also in text understanding. ### Text Understanding Is (Finally) A Reality Text understanding is needed for a wide range of applications: content moderation, topic classification, customer feedback analytics, sentiment analysis, but also search and chatbots. Currently, tackling these applications requires machine learning experts, an intense labeling effort, and the outcome is often disappointing. NuMind is born out of our own frustration with this situation - it was clear that things should be done differently in a post GPT-3 world already, so we decided to do something about it. I (Etienne, CEO) left my position of head of Machine Learning at Wolfram Research, Samuel (CTO) left his previous startup Make.org, and we teamed up to build NuMind - a tool to create custom NLP models efficiently, that can be used by both experts and non-experts. Note the importance of the "custom" aspect here. It is our experience that most NLP tasks are unique - even for something as standard as sentiment analysis - which makes off-the-shelf models not so useful. ### Interactive AI Development The long-term idea of NuMind is that we should teach computers in the same way as we would teach humans. Think of how you would proceed with an intern whose job is to classify your emails for example. You would first give instructions and examples, and then a conversation would start: your intern would ask you questions, and in return you would test your intern to identify and correct their knowledge gaps. We call this approach interactive AI development to stress the importance of the human-AI interaction in order to define/teach a task. Interactive AI development is both natural and extremely efficient, and this paradigm can now be put into action thanks to recent large language models. Of course the road will be long - a lot of R&D is needed - but we started to implement it for text understanding applications and it works! ### Try It Out! We released a private beta a few months ago and have a dozen customers, with use cases such as sentiment analysis, job offer classification, or legal document analysis. We also currently offer a founder-level support to help you succeed your NLP project. So if you think you might have a need for automatic text understanding, do not hesitate to book a demo . Screenshot of the NuMind application showing a text classification project being labelled and trained