# TextUnbox documentation

<img src="docs/textunbox_logo.png" width="250px" />

Welcome to the TextUnbox documentation page.

# What is TextUnbox?

A Software as a Service solution that utilizes AI to offer different services like extraction of **printed** and **handwritten** text from static/non-selectable content, extract text from speech, generate image from text, translate text and so on.

# How does it work?

1. Select an image or an audio file that contains the text you want to extract.
2. Send the image to the TextUnbox service.
3. The recognized text is returned from the service.

The binary content is send to the Azure cloud, where the text is extracted and returned to your device for further use.

> The usage of the service requires a working internet connection in order to verify the license key and process the extraction. A license key can be purchased directly from the [home page](https://textunbox.app/#pricing) or from the TextUnbox product page on [Gumroad](https://keenthinker.gumroad.com/l/textunbox).

# The web applications

Navigate to the [TextUnbox products page](https://textunbox.app/products) and follow the instructions to use the web applications. Following applications are available online at the moment:

- [Text and description extraction ](https://textunbox.app/textunbox)
- [Voice drawing](https://textunbox.app/voicedraw)
- [Text drawing](https://textunbox.app/textdraw)
- [Audio transcription](https://textunbox.app/audiounbox)
- [Translation](https://textunbox.app/texttranslate)
- [Remove image background](https://textunbox.app/products) **under development**!
- [Text to audio](https://textunbox.app/products) **under development**!
- [Chat using ChatGPT](https://textunbox.app/products) **under development**!

# The API

The Textunbox REST API offers different endpoints for extracting text from images and other operations and uses standard HTTP response codes, authentication, and verbs.

- Depending on the endpoint requests with `Content-Type` set to `multipart/form-data` or `application/json` are accepted
- Responses are JSON-encoded (`Content-Type` is set to `application/json`)
- JSON *result* object: ```{
  "success": boolean,
  "result": string,
  "error": string,
  "statusCode": string
}``` 

A detailed API description can be found in the [OpenAPI specification](https://textunbox.app/textunboxapidefinitions.yaml).

## Authentication

You need a valid license key to use the API. A license key can be obtained from the TextUnbox product page on [Gumroad](https://keenthinker.gumroad.com/l/textunbox). 

Define a header with the name `x-textunbox-licensekey` and the **license key** as a value for each API call.

All API requests must be made over **HTTPS**. Calls made over plain HTTP will fail. API requests without authentication will also fail.

The Twitter bot can also be used with the API license key. The same limits apply to the bot as to the API.

## Language

You can specify the language (e.g for the image extraction) using the custom header with the name `x-textunbox-language`. 

Pass the language code string, e.g. `de-DE` or `de`, as a value. Please note, that the different endpoints acceppt language values in different formats. 

## Responses

Every response always has the same structure. A call to an endpoint always returns the *result* JSON object and a standard HTTP response code. The response codes are described in the chapter [Results and Errors](?id=results-and-errors).

The *result* fields have the following meaning:

- *success* is `true` if the extraction was successful; `false` otherwise (examples: monthly usage limit exceeded or license key is invalid or not supported image type)
- *result* the recognised text if success is `true`
- *error* detailed error message if success is `false`; empty otherwise
- *statusCode* the HTTP response status code [List of codes - MDN](https://developer.mozilla.org/en-US/docs/Web/HTTP/Status)

The `ExtractWrittenTextFromImageArea` endpoint extends the *result* object with an additional *meta* attribute, which is a JSON object with the following fields:
- *textWithBoundingBoxItems* is an array of JSON objects with two properties
- - *textWithBoundingBoxItem* is a JSON object with two properties 
- - - *boundingBox* is an array of doubles representing the coordinates of the area framing the recognized text, e.g. `[10.0, 25.0, 56.7, 25.4, 56.5, 50.1, 10.9, 50.0]`; Please see the image below for a visual explanation of the bounding box and the coordinates
- - - *text* is the extracted text which is framed within the bounding box
- - *textIfBoundingBoxWasHit* is a string containing the text that was extracted if the bounding box that was specified encloses an existing recognized bounding box; otherwise if there is no match this property is empty

Example:

<img src="docs/textunbox_boundingbox_explanation.png" />

The above image widht and height is 572x353 pixels. Calling the endpoint with the value `[121.0,189.0,194.0,181.0,196.0,203.0,123.0,211.0]` for the header `x-textunbox-boundingbox` returns the following response:

```JavaScript
{
    "meta": {
        "textWithBoundingBoxItems": [
            {"boundingBox": [52.0,21.0,286.0,21.0,286.0,39.0,52.0,39.0],"text": "bounding box coordinates"},
            {"boundingBox": [47.0,70.0,288.0,70.0,288.0,90.0,47.0,90.0],"text": "[x1,y1,x2,y2,x3,y3,x4,y4]"},
            {"boundingBox": [67.0,156.0,133.0,155.0,134.0,172.0,67.0,173.0],"text": "(x1,y1)" },
            {"boundingBox": [170.0,149.0,236.0,148.0,236.0,169.0,170.0,171.0],"text": "(x2,y2)"},
            {"boundingBox": [437.0,159.0,554.0,159.0,554.0,176.0,437.0,176.0],"text": "bounding box"},
            {"boundingBox": [121.0,189.0,194.0,181.0,196.0,203.0,123.0,211.0],"text": "Hello"},
            {"boundingBox": [74.0,225.0,150.0,224.0,150.0,241.0,74.0,243.0],"text": "(x4, y4)"},
            {"boundingBox": [185.0,218.0,259.0,219.0,259.0,236.0,185.0,236.0],"text": "(x3, y3)"},
            {"boundingBox": [323.0,264.0,428.0,279.0,419.0,324.0,318.0,309.0],"text": "world"}
        ],
        "textIfBoundingBoxWasHit": "Hello"
    },
    "result": "OK",
    "success": true,
    "error": "",
    "statusCode": 200
}
```

## Results and Errors

The TextUnbox API uses conventional HTTP response status codes to indicate the success or failure of a request. 

In general: 
- Codes in the `2xx` range indicate success 
- Codes in the `4xx` range indicate an error that failed given the information provided (e.g. monthly usage limit exceeded or license key is invalid or not supported image type and so on)
- Codes in the `5xx` range indicate an error within the TextUnbox servers

TextUnbox returns one of the following codes:

- `200` Successful extraction operation
- `400` A general error occurred
- `401` Unauthorized; key not present as request header OR key is empty OR key is invalid
- `412` (for text extraction) Image parameter is missing or too many parameters were specified OR Image parameter size must be less than 4MB OR Image dimensions must be between 50x50 and 4200x4200 pixels OR Not supported image format - only JPEG, PNG, GIF and BMP are acceptable OR bounding box header is missing or has an incorrect format
- `412` (for image generation) Request json should not be empty. Example: { prompt: "cyborg flying in space, high resolution", size: "256x256" } OR Prompt is either missing or is empty. OR Size string is either missing or is empty. OR Prompt should not exceed 1000 characters. OR Size string has wrong content. Available image size values are 256x256, 512x512, or 1024x1024.
- `412` (for text from audio extraction) Audio file parameter size must be less than 50MB
- `429` Limit of requests reached OR Subscription request limit reached for the current period; Usage will be available again at the beginning of the new subscription period
- `500` An error occurred while extracting the text

## Endpoints

All endpoints require the POST-HTTP verb. The language is recognized automatically, but it can be also specified explicitely as a request header with the name `x-textunbox-language`. The value should contain the language code. Supported language codes (in alphabetical order) are specified in the endpoint documentation below.

### ExtractPrintedTextFromImage

Extract printed text from an image. 

*parameters*
- `image` (**required**) 
Image file to upload. Please specify the `Content-Type` of the image parameter, e.g. `image/png`.

*returns*

Returns the *result* object

|Supported languages|Code|
|----|----|
|Arabic|ar|
|Bulgarian|bg|
|Czech|cs|
|Danish|da|
|English|en|
|Spanish|es|
|Finnish|fi|
|French|fr|
|German|de|
|Greek|el|
|Hungarian|hu|
|Italian|it|
|Japanese|ja|
|Korean|ko|
|Norwegian|no|
|Dutch|nl|
|Polish|pl|
|Portuguese|pt|
|Romanian|ro|
|Russian|ru|
|Serbian-Cyrillic|sr-Cyrl|
|Serbian-Latin|sr-Latn|
|Swedish|sv|
|Turkish|tr|
|Chinese Simplified|zh-Hans|
|Chinese Traditional|zh-Hant|

### ExtractWrittenTextFromImage

Extract handwritten text from an image. This method is slightly slower than the **ExtractPrintedTextFromImage**, but more accurate and it works also for printed text.

*parameters*
- `image` (**required**) 
Image file to upload. Please specify the `Content-Type` of the image parameter, e.g. `image/png`.

*returns*

Returns the *result* object

|Supported languages|Code|
|----|----|
|German|de|
|English|en|
|Spanish|es|
|French|fr|
|Italian|it|
|Japanese|ja|
|Korean|ko|
|Portuguese|pt|
|Chinese Simplified|zh-Hans|

### ExtractWrittenTextFromImageArea (*beta*)

Extract handwritten text from a specified image area. The recognition area is specified by a bounding box - an array containing exactly 8 elements of type double. 

Example: `[1.0, 2.0, 3.0, 4.0, 5.0, 6.0, 7.0, 8.1]` 

The bounding box is the *border* of the recognized text. Currently it is checked if the recognized coordinates are within a range of values (per value). Currently the *range* for the check is hardcoded to `5` pixels, that are added and substracted from every coordinate position (this parameter will be made dynamic in the future). All values are in pixels (so must be in the image)! 

Example: It was detected that the text `ABC` starts at `X1 = 202`. If for BoundingBox the value `200` is transmitted, it is checked whether `202` lies between `195 (=200-range)` and `205 (=200+range)`, or expressed as a formula `195 <= 202 <= 205`. 

The method returns all recognized texts and their corresponding BoundingBoxes. In general, it is enough to take the BoundingBox for the desired text area and set the value of the parameter `x-textunbox-boundingbox` and subsequent recognitions will check the specified area and extract the new text. 

[Explanation from the Microsoft documentation](https://westus.dev.cognitive.microsoft.com/docs/services/computer-vision-v3-2-preview-1/operations/5d9869604be85dee480c8750): 

> Quadrangle bounding box of a line or word, depending on the parent object, specified as a list of 8 numbers. The coordinates are specified relative to the top-left of the original image. The eight numbers represent the four points, clockwise from the top-left corner relative to the text orientation. For image, the (x, y) coordinates are measured in pixels. For PDF, the (x, y) coordinates are measured in inches.\". 

<img src="docs/textunbox_boundingbox_explanation.png">

The above image widht and height is 572x353 pixels. Calling the endpoint with the value `[121.0,189.0,194.0,181.0,196.0,203.0,123.0,211.0]` for the header `x-textunbox-boundingbox` returns the following response:

```JavaScript
{
    "meta": {
        "textWithBoundingBoxItems": [
            {"boundingBox": [52.0,21.0,286.0,21.0,286.0,39.0,52.0,39.0],"text": "bounding box coordinates"},
            {"boundingBox": [47.0,70.0,288.0,70.0,288.0,90.0,47.0,90.0],"text": "[x1,y1,x2,y2,x3,y3,x4,y4]"},
            {"boundingBox": [67.0,156.0,133.0,155.0,134.0,172.0,67.0,173.0],"text": "(x1,y1)" },
            {"boundingBox": [170.0,149.0,236.0,148.0,236.0,169.0,170.0,171.0],"text": "(x2,y2)"},
            {"boundingBox": [437.0,159.0,554.0,159.0,554.0,176.0,437.0,176.0],"text": "bounding box"},
            {"boundingBox": [121.0,189.0,194.0,181.0,196.0,203.0,123.0,211.0],"text": "Hello"},
            {"boundingBox": [74.0,225.0,150.0,224.0,150.0,241.0,74.0,243.0],"text": "(x4, y4)"},
            {"boundingBox": [185.0,218.0,259.0,219.0,259.0,236.0,185.0,236.0],"text": "(x3, y3)"},
            {"boundingBox": [323.0,264.0,428.0,279.0,419.0,324.0,318.0,309.0],"text": "world"}
        ],
        "textIfBoundingBoxWasHit": "Hello"
    },
    "result": "OK",
    "success": true,
    "error": "",
    "statusCode": 200
}
```

If there was no match for the specified bounding box value, all recognized texts and their respective bounding box are listed and the `textIfBoundingBoxWasHit` property is empty. 

This method is slightly slower than the **ExtractPrintedTextFromImage**, but more accurate and it works also for printed text.

*parameters*
- `image` (**required**) 
Image file to upload. Please specify the `Content-Type` of the image parameter, e.g. `image/png`.

*returns*

Returns the extended *result* object with the additional *meta* attribute (see the above example)

|Supported languages|Code|
|----|----|
|German|de|
|English|en|
|Spanish|es|
|French|fr|
|Italian|it|
|Japanese|ja|
|Korean|ko|
|Portuguese|pt|
|Chinese Simplified|zh-Hans|

### ExtractDescriptionFromImage

Analyzes an image and generate a human-readable short description of it. At the moment, English is the only supported description language. 

*parameters*
- `image` (**required**) 
Image file to upload. Please specify the `Content-Type` of the image parameter, e.g. `image/png`.

*returns*

Returns the *result* object. The `result` parameter contains a description in English language of the image.
### RemoveImageBackground

Enhances images by removing the background, leaving only the foreground object. This is done by dividing the image into multiple segments or regions. The returned image contains only the foreground object and the background is made transparent. 

Left is the original image and right the enhanced image:

<img src="docs/textunbox_removebackgroundimage_example.jpg" width="50%" />

> * Please note that objects that are not in the foreground may not be recognized as part of the foreground. 
> * Background removal works best for categories such as people and animals, buildings and environmental structures, furniture, vehicles, food, text and graphics, and personal belongings.
> * Images with thin and detailed structures, like hair or fur, may show some artifacts when overlaid on backgrounds with strong contrast to the original background.

*parameters*
- `image` (**required**) 
Image file to upload. Please specify the `Content-Type` of the image parameter, e.g. `image/png`.

*returns*

Returns the *result* object. The `result` parameter holds the enhanced image data formatted as Base64 string.


### TranslateText

Translates the input text in the specified destination language. 

The source language is specified in the request body as plain text. The source language is atomatically recognized by the service.

The destination language **must** be defined in the custom header `x-textunbox-language`. 

|Supported languages|Code|
|----|----|
|Arabic|ar|
|Albanian|sk|
|Bulgarian|bg|
|Chinese Simplified|zh-Hans|
|Chinese Traditional|zh-Hant|
|Croation|hr|
|Czech|cs|
|Danish|da|
|Dutch|nl|
|English|en|
|Estonian|et|
|Spanish|es|
|Finnish|fi|
|French|fr|
|Georgian|ka|
|German|de|
|Greek|el|
|Hindi|hi|
|Hungarian|hu|
|Icelandic|is|
|Indonesian|id|
|Italian|it|
|Japanese|ja|
|Korean|ko|
|Latvian|lv|
|Lithuanian|lt|
|Macedonian|mk|
|Norwegian|no|
|Polish|pl|
|Portuguese|pt|
|Romanian|ro|
|Russian|ru|
|Serbian-Cyrillic|sr-Cyrl|
|Serbian-Latin|sr-Latn|
|Slovak|sk|
|Slovenian|sl|
|Spanish|es|
|Swedish|sv|
|Turkish|tr|
|Ukrainian|uk|
|Vietnamese|vi|

### ExtractTextFromAudio

Transcribes audio into text (speech to text). The input format is WAV (16 kHz oder 8 kHz, 16 Bit und Mono-PCM).

The default extraction language is English (`en-US`). 

If your audio input is not in English, please specify the audio input language as the value of the custom header `x-textunbox-language`.

*parameters*
- `audio` (**required**) 
WAV file to upload.

|Supported languages|Code|
|---|---|
|Albanian|sq-al|
|Armenian|hy-am|
|Bulgarian|bg-bg|
|Bosnian|bs-ba|
|Chinese|zh-cn|
|Croatian|hr-hr|
|Czech|cs-cz|
|Danish|da-dk|
|Dutch|nl-nl|
|Georgian|ka-ge|
|German|de-de|
|Greek|el-gr|
|Gujarati|gu-in|
|English|en-us|
|Spanish|es-es|
|Estonian|et-ee|
|Finnish|fi-fi|
|French|fr-fr|
|Irish|ga-ie|
|Hindi|hi-in|
|Hungarian|hu-hu|
|Indonesian|id-id|
|Icelandic|is-is|
|Italian|it-it|
|Japanese|ja-jp|
|Kazakh|kk-kz|
|Korean|ko-kr|
|Lithuanian|lt-lt|
|Macedonian|mk-mk|
|Mongolian|mn-mn|
|Malay|ms-my|
|Norwegian|nb-no|
|Polish|pl-pl|
|Portuguese|pt-pt|
|Romanian|ro-ro|
|Russian|ru-ru|
|Slovak|sk-sk|
|Slovenian|sl-si|
|Serbian|sr-rs|
|Swedish|sv-se|
|Thai|th-th|
|Turkish|tr-tr|
|Ukrainian|uk-ua|
|Uzbek|uz-uz|
|Vietnamese|vi-vn|

### GenerateImageFromPrompt

Generates an image from a text description (prompt) using the [OpenAI DALL.E](https://openai.com/dall-e-3/) engine. Both versions `2` and `3` are supported.
* [OpenAI DALL.E 2](https://openai.com/dall-e-2/)
* [OpenAI DALL.E 3](https://openai.com/dall-e-3/)

The input parameter is a JSON string containing the description and the size of the generated image, specified in the request body with `application/json` as content type.

```
{
  "model": "dall-e-3",
  "prompt": "image description",
  "size": "1024x1024",
  "style": "natural",
  "quality": "standard"
}
```

Parameters and their values:
* `model`: DALL.E engine version. Possible values are `dall-e-2` and `dall-e-3`.
* `prompt`: the description of the image.
* `size`: 
  * `256x256`, `512x512`, `1024x1024` for version 2
  * `1024x1024`, `1792x1024`, `1024x1792` for version 3 
* `style`: `natural` or `vivid`
* `quality`: `standard` or `hd`

All parameters are **mandatory**. If DALL.E version 2 is specified as the model, then the values in `style` and `quality` are not used, since they are supported only from DALL.E version 3. 

OpenAI understands various languages as input prompt. If generation is not working properly with your language, translate the description into English, which is for sure supported. 

The returned generated image data is formatted as Base64 string.   

## Examples

### Postman

The TextUnbox API specification as a Postman collection: 
- [textunboxpostmancollection.json](https://textunbox.app/textunboxpostmancollection.json).


### C#

The easiest way to consume the API with C# is with the [RestSharp](https://restsharp.dev/) library.

```C#
var client = new RestClient();
var request = new RestRequest("https://hello.textunbox.app/api/ExtractPrintedTextFromImage", Method.Post);
// OR
// var request = new RestRequest("https://hello.textunbox.app/api/ExtractWrittenTextFromImage", Method.Post);
request.AddHeader("Content-Type", "multipart/form-data");	
request.AddFile("image", "<path to image>", "<image mime type, for example: image/png>");
request.AlwaysMultipartFormData = true;
request.AddHeader("x-textunbox-licensekey", "<license key>");
//optional
//request.AddHeader("x-textunbox-language", "<language code>");
var response = await client.ExecuteAsync(request);
Console.WriteLine(response.Content);
```

### JavaScript (client side)

The `file` parameter in the following code examples is the value of the input field of type `file` used to select the file (`$('#imagetoscan')[0].files[0]`). The `apikey` parameter is a string holding the license key value.

```HTML
<input 
  type="file" 
  id="imagetoscan" 
  name="imagetoscan" 
  accept="image/png, image/jpeg">
```
Consuming the API using [Fetch API](https://developer.mozilla.org/en-US/docs/Web/API/Fetch_API) looks like this:

```JavaScript
var headers = new Headers();
headers.append("x-textunbox-licensekey", apikey);
//optional
//headers.append("x-textunbox-language", "de");

var formdata = new FormData();
formdata.append("value", file, file.name);

var requestOptions = {
  method: 'POST',
  headers: headers,
  body: formdata,
  redirect: 'follow'
};

fetch("https://hello.textunbox.app/api/ExtractPrintedTextFromImage", requestOptions)
  .then(response => response.text())
  .then(result => console.log(result))
  .catch(error => console.log('error', error));
```

Consuming the API using [jQuery ajax API](https://api.jquery.com/jquery.ajax/)

```JavaScript
var form = new FormData();
form.append("value", file, file.name);

var settings = {
  "url": "https://hello.textunbox.app/api/ExtractWrittenTextFromImage",
  "method": "POST",
  "timeout": 0,
  "headers": {
    "x-textunbox-licensekey": apikey,
    "accept": "application/json",
  },
  "processData": false,
  "mimeType": "multipart/form-data",
  "contentType": false,
  "dataType": "json",
  "data": form
};

$.ajax(settings).done(function (response) {
  console.log(response);
});
```

### JavaScript (server side)

Consuming the API using NodeJs, axios and the [form-data](https://www.npmjs.com/package/form-data) npm package

```JavaScript
const FormData = require("form-data");

const fs = require("fs");
const path = require("path");
const form = new FormData();
const image = "image.png";

const rs = fs.createReadStream(path.join(__dirname, image));
form.append('value', rs, image);

const apiUrl = 'https://hello.textunbox.app/api/ExtractPrintedTextFromImage';

axios.post(apiUrl, form, { 
    headers: {
        "x-textunbox-licensekey": apikey,
        "accept": "application/json",
        ...form.getHeaders()
    }
}).then(function (response) {
    console.log(response);
}).catch(function (error) {
    console.log(error);
});
```

## OpenAPI specification

The TextUnbox API specification in OpenAPI (Swagger) format:

- [YAML](https://textunbox.app/textunboxapidefinitions.yaml)
- [JSON](https://textunbox.app/textunboxapidefinitions.json)
