Unstructured data analysis

Important

This feature is available through the Early Adopter Program (EAP). To evaluate this feature, contact Diligent Support.

Unstructured data analysis enables you to work directly with PDFs in your datasets. You can extract clean text from PDFs and convert the content into structured tables for further analysis.

It helps you:

  • Analyze large volumes of document content more efficiently.

  • Identify patterns, themes, and useful insights from unstructured data.

  • Highlight potential anomalies or risks for faster review.

  • Reduce the manual effort required to process complex documents.

  • Accelerate workflows such as bulk invoice processing, investigations, compliance and audit reviews, research, and risk analysis.

Key capabilities

  • Single and bulk upload Upload one or many documents directly into a dataset.

  • Immediate extraction on upload The AI-powered extraction enables a quick preview of the data.

  • Side‑by‑side viewer View the original PDF page next to the extracted text for direct comparison and review.

  • Easy verification Highlight text in the Markdown and see the corresponding region highlighted in the PDF, making it easy to verify extraction accuracy and context.

  • Manual corrections Edit the extracted text directly to fix AI errors; changes are tracked for audit purposes.

Upload a document

  1. On the ACL AI Studio home page, from the list of datasets, select the dataset for which you want to upload a PDF.

  2. On the Dataset summary page, from the left navigation, select Documents.

  3. In the Documents panel, select Upload new.

  4. In the Upload documents panel, add a PDF by browsing your local system or dragging it into the designated area.

    You can upload up to 10 PDFs at a time.

  5. Select Add.

    The Documents panel displays the list of documents added.

Understanding confidence score

A confidence score indicates how reliable the extracted content is for a processed page. It helps you identify pages that may need closer review before the document is used in analysis.

On a 0% to 100% scale, the ratings are as follows:

  • High (75% and above) Indicates high confidence in the extracted text.

  • Medium (40-74%) Indicates that some extracted values should be reviewed.

  • Low (less than 40%) Indicates that the extraction may contain inaccuracies and should be verified carefully or re-extracted.

Validate an extracted document

Validate extracted text to confirm that the output matches the source PDF and is ready for analysis.

  1. From the list of documents under Ready to preview, select the document you want to validate.

    The screen that appears displays the original document and the extracted text side-by-side.

  2. Review the original document in the Document viewer pane.

  3. In the Extracted text pane, compare the extracted text with the original document (page-by-page) to ensure that key information is extracted correctly.

  4. Review the confidence level displayed for the extraction.

  5. If necessary, use the edit icon () to modify the extracted text.

  6. If the extracted text is inaccurate or incomplete, select Re-extract to extract that page again.

  7. Select Validate.

  8. In the confirmation dialog, select Validate.

Generate structured output from a validated document

  1. From the list of documents under Validated, select the document that you want to convert into a structured output.

  2. On this screen that appears, select an AI-generated suggestion or start a chat.

    Your chat history is automatically saved with the document, so any user with access to the document can view the chat history and continue making requests from where the previous conversation left off.

  3. Optional If you want to start a new conversation, select Reset chat.

    This clears the current chat context and lets you start fresh.

  4. Review the sample preview and refine if required.

  5. Select Expand to view the samples on the left pane.

  6. On the left pane, select View all records.

    The entire table is displayed.

  7. Optional Select the View link in the Reference column of a record to view the source page from which the data is extracted.

    The relevant text in the PDF is highlighted, indicating the specific source content corresponding to that row.

  8. Select Save as Source table.

  9. Add a name to the table and select Save.

    The Source tables section in the Tables list displays the newly added table.

    You can now use this table for analysis. For more information, see Analyze your data.

View the details of an extracted document

  1. From the list of documents, hover on the document for which you want to view the details.

  2. Select the three dots that appear to the right of the document name, and select View details.

    In the side panel that appears you can view the following details:

    • Size of the file.

    • The date and time when the file was uploaded.

    • User who uploaded the file.

    • The date and time when the file was validated.

    • User who validated the file.

Delete a document

  1. From the list of documents, hover on the document that you want to delete.

  2. Select the three dots that appear to the right of the document name, and select Delete.

  3. In the confirmation box, select Yes.