API WORKFLOWS

Document Processing APIs

Choose a document API by the input you have and the output your application needs. Use PDF generation for reports, PDF text extraction for document review, and image OCR for photographed or scanned text.

A practical integration workflow

  1. Identify the source format and check file size and page limits.
  2. Send a request to the matching endpoint with your API key.
  3. Check HTTP and application status, then save the returned file or text.

Choose an API and try an example

Each example uses the normal API endpoint. Replace YOUR_APPKEY with your own key. Use the separate Demo link to view fixed sample data without a subscription.

HTML to PDF

Submit HTML for a report or a public URL, then save the PDF returned by the API.

Input and coverage: HTML inputs support up to 5 MiB; URL inputs up to 2048 characters. Keep access to your input and save the returned PDF.

Example request

Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.

curl --fail-with-body 'https://api.gugudata.io/v1/imagerecognition/html2pdf?appkey=YOUR_APPKEY' \
  --request POST \
  --header 'Content-Type: application/json' \
  --data '{"type": "html", "content": "<html><body><h1>Quarterly report</h1><p>Prepared for review.</p></body></html>", "landscape": 0}'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "pdfPath": "<temporary PDF URL>"
  }
}

PDF to Text

Extract text from a PDF for review, search, or a downstream document workflow.

Input and coverage: Maximum 20 MiB and 50 pages. Supply either file or file_url. Supported lang values are en, zh-CN, zh-TW, en+zh-CN, and en+zh-TW.

Example request

Replace YOUR_APPKEY with your own API key. This request uses the normal API endpoint.

curl --fail-with-body 'https://api.gugudata.io/v1/imagerecognition/pdf2text?appkey=YOUR_APPKEY' \
  --form 'file_url=https://gugudata.io/assets/demo/pdf2text-demo.pdf' \
  --form 'lang=en'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "resultText": [
      "GuGuData PDF OCR demo"
    ],
    "text": "GuGuData PDF OCR demo\n\nSecond page document workflow"
  }
}

Image OCR

Read text from an image using a multipart file upload.

Input and coverage: JPEG, PNG, WebP, TIFF, or BMP input, maximum 10 MiB. Recognition quality depends on resolution, contrast, and layout.

Example request

Download the public sample image first. Replace YOUR_APPKEY with your own key before sending the API request.

curl --fail --output gugudata-logo.webp 'https://gugudata.io/assets/svg/logos/logo.webp'

curl --fail-with-body 'https://api.gugudata.io/v1/imagerecognition/ocr?appkey=YOUR_APPKEY' \
  --form '[email protected]'
View sample Demo response

This is a fixed Demo sample, shortened for readability. It is not the response to your example request. Your results depend on the input and dataset. File URLs are temporary and have been removed.

{
  "dataStatus": {
    "statusCode": 200,
    "status": "SUCCESS",
    "statusDescription": "successfully",
    "dataTotalCount": 1
  },
  "data": {
    "resultText": [
      "GUGUDATA"
    ],
    "text": "GUGUDATA\nDATA PAY LESS"
  }
}

Limits and things to check

PDF generation and OCR solve different tasks. OCR quality depends on the input; test representative documents before relying on extracted text. These APIs do not verify the accuracy or legal meaning of a document.

Normal paid API subscriptions include unlimited total requests during the active paid term, with up to 5 requests per second. Free trials have a limited allowance: sign in and check Trial Keys for the balance. Read each product page for pricing, renewal, and cancellation details.