Dataset Builder

Turn documents, images and recordings into AI-ready text, right on your Mac

Drop in one file or an entire archive. The app reads, transcribes, tidies, splits and source-marks the content, without uploading any of it.

The result is tidy text files AI tools can read, where every passage points back to its source. No account, no sign-in, no cloud service.

From NOK 1,790 excl. VAT · One machine, one year

The app in brief

From one PDF to the whole archive

One PDF, a photo of a document, a spreadsheet, a meeting recording or a folder tree holding thousands of files. The same workflow at any volume, and ZIP archives are unpacked along the way.

Speech becomes text

Meeting recordings, interviews and voice notes become timecoded text you can search and quote. Every passage points to the right second in the recording.

Your content is not uploaded

The app has no network entitlement and processes everything locally. macOS enforces it, and your IT department can verify the entitlements themselves.

What must stay out, stays out

List names, code words and set phrases, and they are removed from all exported content, with the app checking afterwards that they are actually gone.

Pre-split for your AI tool

The content is split along its natural boundaries, into files the tool can take in at once, and very large blocks that do have to be divided are clearly marked.

Every answer can be checked

Every part of the result states its document, page, sheet, slide or timecode, so AI answers can be looked up against the original.

What can you drop in?

Thousands of files in a single run, or just the one. Everything is processed locally and comes out as tidy text.

  • Documents PDF, Word, PowerPoint, Markdown, HTML
  • Spreadsheets Excel, CSV: sheets and rows kept together
  • Images containing text Screenshots, photos of documents, scanned pages
  • Voice recordings Apple Voice Memos, interviews, notes
  • Meeting recordings Recordings from Teams and other meeting tools
  • Entire folder trees Point at a root: ZIP files are unpacked along the way
Berigo Dataset Builder Reads, transcribes, tidies and splits Everything happens locally
A dataset ready for AI Text files where every passage points back to its place in the original: page, sheet, slide or timecode in the recording.

Language support, in three parts

Documents, spreadsheets and presentations are read directly, with no recognition involved. Every language and script works.

Text in images and scanned pages is recognised in around 30 languages by macOS' built-in text recognition, among them English, Norwegian bokmål and nynorsk, Swedish, Danish, German, French, Spanish, Arabic, Chinese and Japanese.

Speech in audio and video recordings is transcribed in English, Norwegian bokmål, nynorsk, Swedish and Danish. Through the speech recognition in macOS, German, French, Spanish, Italian, Portuguese, Japanese, Korean and Chinese, both Mandarin and Cantonese, are supported as well. The language is detected automatically per recording.

What do you get out?

The result is called a dataset: tidy text files where your content has been extracted, put in order and cleared of noise, ready to hand to an AI tool. Concretely, you get:

  • Text files at the right size. An AI tool can only read a limited amount of text at a time, called the context window. The files are split accordingly, along the content’s natural boundaries. If a very large table still has to be divided, the continuation is clearly marked with the same source mark.
  • A source mark on every passage. Document number, file path and a precise location: page, sheet and rows, slide, timecode or heading.
  • Timecodes in the transcripts. Every passage from a recording points to the right second in the original.
  • An index register. One line per source file: what was read, what was left out and why, measured quality and warnings, with a digital fingerprint per output file that exposes any later change.
  • Formats you choose yourself. Markdown as standard, with plain text and JSON Lines in addition.

The difference between this and a random folder of text files is that the result can be used: well-split material gives better answers, and the source marks make the answers checkable.

Speech becomes text

A video meeting often exists only as a recording. An interview sits in an audio file. A voice note was recorded in the car. Drop the recording in, and get back text where every passage carries a timecode that leads you to the right place in the original.

Transcribe recordings in English, Norwegian bokmål, nynorsk, Swedish and Danish. Recognition for Norwegian, Swedish and Danish is built into the app itself and works without a network connection. Danish, which many tools struggle with, is transcribed with the same built-in quality; Danish transcripts are delivered in normalised form, in lower case and without punctuation. Through the speech recognition built into macOS, the app additionally supports German, French, Spanish, Italian, Portuguese, Japanese, Korean and Chinese, both Mandarin and Cantonese.

The app accepts recordings from Apple Voice Memos directly, along with common audio and video formats such as MP3, M4A, WAV, MP4 and MOV. The language of each recording is detected automatically. The app processes recording files and does not record meetings itself, and a recording it cannot render reliably is set aside with an explanation rather than guessed at.

How it works

The screenshots below show the Norwegian interface. The app itself runs in English too: choose your language under Settings.

1 Choose the material. Drop in one file, or point at a folder. The app gets access to what you point at, and nothing else. Sample data ships with the app, so you can try the whole workflow without touching real documents.
2 See what the app found. One overview of every file, with duplicates, scanned pages, encrypted files, measured text quality and estimated volume.
3 Shape the datasets. Each top-level folder becomes a dataset. Merge, split, hand-pick individual files or leave them out, so an AI tool can get one subject area without getting the rest.
4 Set the rules. Decide how much structure is kept, and list words and phrases that must never come along. Save the rule set as a profile for next time.
5 Check before you run. See real content before and after, formatted exactly as the export will be.
6 Export. Tidy text files per dataset, pre-split, with the index register and a receipt showing what actually happened.

Answers you can stand behind

An AI answer without a source is an assertion. Ask the tool to state document numbers in its answers, and every claim can be looked up: the right document, the right page, the right second. The index register and the fingerprints mean every claim can be looked up and every change detected. That holds up for you, for your auditor and for whoever disagrees.

Your content stays with you

The app has no network entitlement and does not upload documents, audio or exports. Processing happens locally, and macOS enforces the restriction rather than us promising it. Source folders are only ever read, never changed, and access can be withdrawn at any time.

Where the datasets are used afterwards is your decision. The details, including the few exceptions that exist, are set out in the technical review at the bottom of this page.

Use the datasets wherever you want

The app makes the dataset and stops there. You choose the AI tool: paste the content into Claude, ChatGPT or another tool your organisation has approved, or run an open model locally on a suitable Mac, entirely offline.

The user guide is built into the app and covers both routes, with a ready-made prompt that tells the tool to answer only from the material and to state a document number in every answer. The app itself never sends the dataset anywhere.

What it costs

One fixed price per machine per year, and volume pricing for more machines.

One machine, one year

NOK 1,790 per machine per year · excl. VAT

  • The licence covers one machine for twelve months from installation, and can be renewed.
  • All updates during the period are included in the price.
  • No licence server and no sign-in: the app has no network access. The expiry date is checked on your own machine, and nothing is sent to us.

More machines

Price on request

  • The price is set by the number of machines and how the app is to be deployed.
  • An installer package for rolling out to many machines, and a named contact with us.
  • Tell us briefly what you need and we will come back with a firm quote.
Request pricing

All prices are stated without value added tax. VAT is added at the applicable rate.

The app requires a Mac with Apple silicon and macOS 26 or later, runs on Mac only, and its interface is available in Norwegian and English.

For organisations and IT

Deployment is a signed DMG for individual machines, or a .pkg through tools such as Jamf, Kandji and Mosyle. If you use Intune, a signed package is supplied on request. There is no licence server to operate and no accounts to administer.

The entitlement list can be verified technically by your own IT, and every run leaves an index register that can be reviewed line by line. The details sit in the technical review below.

Request pricing

Technical information for IT and security

The short version: the app runs in the macOS sandbox with no network entitlement, collects no usage data, never updates itself, and treats every source file as untrusted. The details follow, point by point.

Sandbox and entitlements

The entitlement list in the finished program has exactly three items: run in the sandbox, read and write files you yourself select through the standard open panel, and remember those folders for next time. A network entitlement does not exist in the program, and a missing entitlement cannot be switched on: it has to be added, and the program built and released again. The app gives you the command that reads the entitlements straight out of the installed program, under Settings and Privacy, so you see what macOS actually enforces.

One qualification belongs here: the file entitlement covers both reading and writing in the same item. That the app only reads the source folder and never writes to it is the app’s own behaviour, not something the operating system enforces. Confirming it takes a review of the code.

No collection and no self-updating

The app does not record how it is used and sends neither usage statistics nor error reports. This is not a setting left off: those functions are not in the program. It builds on code that ships with macOS, plus one pinned, statically linked open-source speech recognition component that never fetches anything at runtime, the built-in speech models, and open-licensed typefaces and icons. Everything is locked with checksums in the build process. There is no automatic updating either, so new versions are installed manually or via MDM.

Speech recognition and language add-ons

Recognition for Norwegian, Swedish and Danish ships with the app itself and works without a network connection from the first launch. For certain other languages the app may ask macOS to fetch a language add-on: the download is performed by the operating system outside the app and covers the add-on only, never audio or text from your material. The download is attempted once per language per app session, and is cut off after 60 seconds; and if it fails the media file is set aside with an explanation. A machine used for this may therefore contact Apple the one time such an add-on is missing. That belongs in your assessment.

Files, logs and retention

Output goes into a folder kept out of the search index macOS builds across the machine, until you choose otherwise. There are two local logs: the system log and a per-run diagnostics file. They hold timestamps, file names, paths and outcomes, never text from the documents, and retention defaults to 90 days. Note that file names alone are often sensitive in board work and disputes; set the retention period with that in mind. Passwords for encrypted PDFs are held in memory only for the length of the job. Note also that the guarantee for excluded words covers content: file and folder names are kept for traceability, and the index register warns when a name contains an excluded phrase.

Source files are treated as untrusted

A document can be built to exploit the program that opens it, so no source file is trusted. ZIP archives are unpacked within fixed limits and checked against the archive’s own checksum, encrypted files without a supplied password and unknown formats are refused rather than misread, shortcuts to files elsewhere are never followed, and HTML is read with network fetching switched off, so a purpose-built file cannot pull anything in from outside.

Boundaries the app cannot hold

Three routes sit outside the app’s control, and they belong in an assessment. The clipboard can be read by every program on the machine, and with Handoff on, its contents can turn up on the user’s other Apple devices. Backup and sync copy the datasets off the machine like any other file. And the datasets themselves are as sensitive as the originals: where they live should be somebody’s deliberate decision.

Signing and checksums

The DMG and ZIP edition is signed with a developer certificate issued by Apple and inspected by Apple through notarisation, and the downloads come with a checksum the recipient can recompute and compare. Output is also deterministic: the same source material produces byte-identical files. The installer package for central deployment is not signed the same way today. Jamf, Kandji and Mosyle will install it regardless, while Intune requires a signed package: ask for a signed package before deploying that way. The entitlement list can be verified for every version you adopt, but the checking is yours, not automatic.

Drop in one file, or the whole archive

Get back tidy, source-marked text ready for the AI tool of your choice, without your content being uploaded. Requires a Mac with Apple silicon and macOS 26 or later.