Everyone has retyped text they could already see.

The long error code in a dialog box. Figures in a document someone sent as an image. A command shown in a video. The scanned half of a PDF. Text in an admin console that refuses to be selected. A configuration value inside a screenshot posted to Slack.

The characters are right there. You can read them. You just can’t copy them. So you retype — and you get one character wrong. You search the error code, get zero results, and burn five minutes before spotting the typo. That particular kind of wasted effort is quietly demoralizing.

MacSide AI reads the text the moment the screenshot is taken. Capture, click, paste.

It’s already been read

Screenshots land in the clipboard panel, where a "detected text" button copies out everything inside the image.

The sequence in the video:

Capture a region of the screen. The screenshot lands in the clipboard panel as an image.

That entry carries a detected-text button. Press it, a confirmation appears, and the text from inside the image is on your clipboard.

Then paste. In the video it goes straight into the AI chat in the side panel with ⌘V.

The important part is what’s missing: there’s no “run OCR” step. The reading already happened at capture time. The button just hands you the result when you want it. If you never want it, it stays an ordinary screenshot.

How this differs from the built-in screenshot

macOS does have text recognition. Open an image in Photos or Preview and you can often select text. For some situations that’s genuinely enough.

But the built-in path looks like this: ⌘Shift+4 to capture. It saves to the desktop. Find it in Finder. Open it in Preview. Test whether the text is selectable. Drag to select. Copy. Close Preview. Delete the file later.

The difference is that there’s an app-opening step after the capture — plus one more file on your desktop. That’s why anyone who screenshots regularly ends up with a desktop full of Screenshot 2026-08-04 at 14.32.11.png.

In MacSide AI, captures go into the clipboard panel instead of scattering across the desktop, and the already-extracted text is one button away inside that panel.

For the specific goal of “get the text out,” the number of apps you open is zero.

Pages too long for the screen, in one image

The other thing built-in screenshots handle badly: tall pages.

Terms of service, long articles, Slack history, a full settings list, a chat thread. Anything you have to scroll to see becomes three or four separate captures — and whoever receives them has to view three files in order.

MacSide AI does scrolling capture, so content that doesn’t fit on screen becomes a single image.

That matters most when you’re recording something:

  • Preserving a terms page before and after a change
  • Handing a full error log to a developer
  • Capturing every field of a settings screen for a runbook
  • Saving the relevant section of an article for a citation

And that combined image is, of course, also readable. So you can lift the full text of a long page that wouldn’t let you select it. A web page you can usually just select; an app window or an embedded table is often another story.

Feed the extracted text straight to AI

At the end of the video the copied text goes into the AI chat in the side panel. That’s not a demo flourish — it’s the main use case.

Extracted text is almost always input to the next thing:

  • An error message → paste it and ask what’s causing it
  • A foreign-language image → paste it and translate
  • A dense reference screenshot → extract and reorganize into notes
  • A table capture → extract and reformat
  • Invoice figures → extract and total

What makes this fast is that OCR and AI chat live in the same side panel. There’s nothing to carry between apps. Capture, read, paste, ask — all at the same screen edge.

Try it once during an actual debugging session and the difference is obvious. Capture the error dialog. Copy the detected text. Paste to AI. Get a direction. You never opened a terminal or a browser tab.

Where it earns its keep

  • Error codes and logs — pull them accurately out of dialogs that won’t let you copy
  • Documents sent as images — get figures out of a screenshot-of-a-PDF
  • Video and streams — pause, capture, and lift the command or config value
  • Paper documents — photograph and extract
  • Unselectable web pages — take just the part you need to quote
  • Other people’s screenshots — use a config value from a Slack image without retyping it

That last one comes up more than you’d guess. Text inside a screenshot someone else took is text you can never select. Retyping it is a small tax paid somewhere in every team, every day.

Frequently asked questions

Do I need to run OCR on each capture?

No. Recognition happens at capture time; the detected-text button just retrieves the result when you need it. Ignore it and the capture behaves like any screenshot.

Which languages does it read?

English and Japanese are both supported, including screens that mix them.

What about pages that don’t fit on screen?

Scrolling capture combines them into a single image, and that combined image is readable too.

Do screenshots pile up on my desktop?

No — captures live in the clipboard panel, so your desktop doesn’t fill with files. You save only the ones you actually want.

How is this different from macOS text recognition?

The recognition itself overlaps. The difference is the number of steps to get there: the built-in path means saving a file and reopening it in another app, while this is one button in the flow you’re already in.

Retyping adds nothing to the result

Retyping text has no creative content whatsoever. You’re re-entering characters you can already read, and if you mistype one, the entire effort was wasted.

It still happens every day, everywhere — for no reason other than the text happening to exist in a form that can’t be copied.

OCR changes that premise. Text on screen is copyable. Text inside a screenshot is copyable. Text inside someone else’s screenshot is copyable.

Capture, click, paste. Once it’s that short, a related compromise disappears too — the one where you paraphrase a value instead of quoting it exactly because retyping was annoying. Exact strings stay exact on the way to wherever they’re going.