How to chat with a PDF online in Whizi

Upload a PDF to Whizi, ask questions across its full content, compare answers between Claude and Gemini, and export the summary in seconds.

First, confirm it is reading your document

The most dangerous failure in document chat is not a wrong answer, it is a confident answer drawn from training data instead of from your file. It happens most with documents that resemble something common: a standard commercial lease, a well known regulation, a widely discussed paper. The model knows the genre well enough to produce a plausible answer without consulting your specific text, and nothing in the output signals that it did so.

One habit removes this. Before asking your real question, ask a locating question.

Quote the passage in the attached document that discusses [topic], and give the page or section it appears in. If it is not in the document, say so.

If it returns a real quote, it is reading your file. If it produces a generic paraphrase or cannot find something you know is there, the extraction has failed and every subsequent answer is suspect. That takes ten seconds and it is the difference between a tool and a liability.

Carry the habit into the real questions too: Quote the exact text that supports your answer should appear in any prompt whose answer will be relied on.

Choosing a model for the document

DocumentModelWhy
100+ page reports, filings, whole contract setsGeminiLargest context window, so the document is genuinely present rather than chunked
Contracts, policy, anything where qualification mattersClaudeBest at nuance and at telling you what a document does not say
Invoices, forms, short structured filesGPTFast, and the most reliable at strict extraction into tables or JSON
Papers with figures, scanned pages, diagramsAny multimodal modelReads the page as an image where text extraction falls short

The context window point is worth spelling out. When a document exceeds what a model can hold, it has to be processed in pieces, and questions that require the whole document at once (does anything here contradict clause 14, is this term defined anywhere, which of these three sections is inconsistent) become unreliable. That is the specific reason to reach for the large-context model on long files rather than a general preference.

You can switch models without re-uploading, so read with one and extract with another against the same file. See switching models mid-conversation.

Prompts by document type

Contracts and agreements

From the attached agreement, extract into a table: term and renewal, notice period for termination, payment terms, liability cap, indemnities, governing law, assignment, and any clause that survives termination. For each, give the clause number and quote the operative sentence. Then list separately the obligations that fall on us rather than on them.

Follow up with the question that actually matters: What is in this agreement that is unusual compared to standard terms for this type of contract, and what is conspicuously absent? Absence is where the risk usually lives, and it is the thing a keyword search can never find.

Research papers

For the attached paper, return: the research question, the study design, the sample and population, the primary outcome, the headline result with its effect size, the stated limitations, and the funding source. Then tell me the three claims in this paper that would need independent verification before I cite them.

Long reports and filings

Summarize the attached report in 10 bullets ordered by importance rather than by document order. Then list: the three numbers a decision maker would care about with their page references, anything the report presents as fact without a source, and any place where the summary or executive section disagrees with the detail later in the document.

That last check finds real problems surprisingly often. Executive summaries are written early and edited less than the sections they summarise.

Comparing two documents

Compare contract-a.pdf and contract-b.pdf clause by clause. Return a table of every substantive difference: topic, what A says, what B says, and which favours us. Ignore formatting and numbering differences. List separately anything present in one and missing from the other.

Turning it into something you can use

Rewrite the findings for an executive audience in 150 words. Lead with the decision required. No jargon that is not defined. Flag anything I should verify before circulating this.

Troubleshooting

"I cannot see a document." Confirm the file finished uploading, then reference it explicitly by name in your message. On very long threads, re-attaching or starting a fresh chat with just the file is faster than arguing about it.

Scanned PDF, garbled or missing text. The document is an image and OCR is doing the work. Whizi handles OCR automatically on most plans and you can also ask explicitly for OCR before answering. Accuracy drops with low resolution, unusual fonts, handwriting, and dense tables. Because a misread digit is invisible in a fluent answer, verify every figure from a scanned document against the original page.

Tables come out wrong. PDF tables are notoriously badly structured underneath. Ask for the table to be reproduced verbatim first, check it against the page, and only then ask for analysis. For heavy table work, a multimodal model reading the page as an image is often more reliable than text extraction.

It missed something you know is in there. Ask the locating question. If the model cannot quote a passage you can see, the problem is extraction rather than reasoning. Try another model, or upload just the relevant pages.

The file is too large. Use the large-context model, or split by purpose rather than by page count. Uploading the three sections you actually need beats uploading 400 pages and hoping.

Answers get vaguer as the conversation goes on. Long threads accumulate context that competes with the document. Start a fresh chat with the file and a concise statement of what you need.

Verify before you rely on it

Three checks, in order of importance.

Require quotes on anything consequential. A quoted passage is checkable in seconds; a paraphrase is not. This single practice eliminates most of the risk in document work.

Cross-check with a second model. Ask the same question of a different model against the same file. Agreement is meaningful evidence; disagreement tells you exactly where to look. Whizi's side-by-side comparison exists for this.

Ask what it is uncertain about. Which of your answers above are you least confident in, and what in the document is ambiguous? Models are imperfect at self-assessment, but the ambiguities they surface are usually genuine ambiguities in the source, which is useful information about the document itself.

And the standing rule for anything that carries consequences, meaning legal, financial, medical, or compliance material: the model finds the passage, you make the judgment. It is a reading assistant, not an adviser, and the accountability for a decision made on its output is entirely yours.

Workflow checklist
  • Ask a locating question first to confirm the model is reading your file
  • Require an exact quote for every consequential answer
  • Use the large-context model for long documents so nothing is chunked away
  • Ask what is conspicuously absent, not just what the document says
  • Verify figures from scanned documents against the original page
  • Reproduce tables verbatim and check them before analysing them
  • Cross-check anything important against a second model
Common questions

Frequently asked questions

What is the maximum PDF size?

Limits depend on your plan and the model, and the real constraint is the model context window rather than a file size cap. Gemini in Whizi supports the largest documents, which is why it is the recommendation for long reports, filings, and multi-contract sets. Above that, upload the sections you actually need rather than the entire document, since targeted context usually produces a sharper answer as well.

Does Whizi store my PDFs?

Files are stored securely in your workspace, are not used to train models, and can be deleted at any time. Deleting a file does not remove what was already discussed about it in the conversation, so if the document is sensitive enough to warrant deletion, delete the chat too. Each provider’s data handling policy is available for review before you enable that model.

Can I chat with multiple PDFs at once?

Yes, and comparison across documents is where this is most valuable. Upload several files to the same thread and reference them by name in your prompts, otherwise the model chooses which one to answer from without telling you. Clause-by-clause contract comparison, consolidating findings across reports, and finding contradictions between documents all need everything present at once, which is another reason to use the large-context model.

Does it work with scanned PDFs?

Yes. Scanned files are processed with OCR automatically on most plans, and you can ask the model to OCR and then answer. Accuracy is good on clean scans and degrades with low resolution, unusual fonts, handwriting, and dense tables. Since a misread character produces a wrong number inside an otherwise fluent answer, check any figure taken from a scanned document against the original page.

How do I know the answer actually came from my document?

Require a quote. Ask the model to quote the exact passage supporting its answer along with the page or section, and to say plainly if the information is not in the document. Answers drawn from training data rather than from your file are the most dangerous failure in document chat, precisely because they are plausible for common document types, and demanding a verifiable quote is the only reliable way to tell the difference.