8/29/2026

FAQ

Is my manuscript used to train models?

ProtoGloss itself does not train on customer documents. Your text is sent to the configured model backend during translation, so that backend's own retention and training policy governs that hop. A precise answer depends on the deployment, and if you are evaluating this for a publishing house you should ask us for it explicitly rather than accept a general reassurance. Parsing happens before any model call.

What happens to my files?

Uploads land in a quarantine area, are validated by size and file signature, and are only then promoted. Full parsing runs in a restricted workload that holds no model-provider credentials and has limited network egress. The command line tool defaults to thirty day retention with explicit preview-then-apply deletion.

We do not claim to scan uploads for malware. The interface for it exists but ships as a permissive stub until a vendor is chosen, and telling you otherwise would be untrue.

Can it handle a scanned book?

No, and it says so rather than guessing. Scanned or empty PDFs are rejected with a clear message. There is no OCR in the product, and no image translation. If your source only exists as scans, ProtoGloss is not the right tool until that changes.

Will the output look like my original file?

No. The export is a clean semantic DOCX: headings and lists recreated, correct text direction for the target language, translated content in order. It is built to enter a production pipeline rather than to imitate the source layout. For most publishers that is the more useful of the two, but it is a deliberate choice and you should know it before you start.

How much does it cost?

Every job shows an estimated cost and time before it runs, so you are never spending blind. Commercially the shape is a free evaluation on your own sample, then a per-project fee, then pay as you go. There are no published price tiers, so the honest answer to what a specific book costs is that we quote it after seeing the material.

How is quality measured?

Runs are scored on fidelity, naturalness and idiomaticity, consistency, and readability, with edit load reported separately because edit load is what actually costs you money. There is an in-product quality report, an issue and score panel per segment, and a sampled model-based judge on the literary presets.

We are still building the public evidence base for this. There is no benchmark number we are willing to quote yet, and we would rather say that than invent one.

Which languages does it support?

ProtoGloss is built and evaluated on English and Persian long-form literary work. That is the pair with in-house bilingual evaluation behind it. The engine accepts any target language the configured model supports, and Spanish, German, French, Portuguese, Arabic and Persian are available in the workspace today, with character name mining shipping language packs for those. Quality outside the evaluated pair depends on the model and on your source material.

Can I integrate it?

Yes. There is a partner API with workspace-scoped keys and signed webhooks, aimed at translation-as-a-service teams and integrators. It is sequenced after the hosted product, so it is available to partners through a conversation rather than a self-serve developer signup.

How do I get access?

Access is by invitation. Request one and we will start with a free evaluation on a sample of your own material.