Configuration
NeoRecall reads ~/.neorecall/.env and process environment variables. See the commented .env.example for the complete list.
Essential server valuesβ
| Variable | Purpose | Default |
|---|---|---|
NEORECALL_HOST | Listen address | 127.0.0.1 |
NEORECALL_PORT | HTTP port | 4500 |
NEORECALL_TRUST_PROXY | Trust one reverse-proxy hop | false |
MAX_UPLOAD_BYTES | Maximum live chunk upload | 33554432 |
NEORECALL_SAME_DEVICE_COVERAGE_RATIO | Skip transcription when this fraction of a chunk's timeline already exists on another source of the same device (live take plus later device-file import) | 0.5 |
NEORECALL_CONTEXT_MAX_FILE_BYTES | Maximum original context-file upload | 33554432 |
NEORECALL_CONTEXT_MAX_ITEMS | Maximum context items per recording or memory | 200 |
NEORECALL_REQUIRE_VECTOR | Fail without the tested sqlite-vec extension | production: true |
NEORECALL_PLAUD_CLIENT_ID / NEORECALL_PLAUD_CLIENT_SECRET | Partner credentials so iOS and Android can bind Plaud Note Pro / NotePin S over BLE | unset (pairing hidden) |
External inference providersβ
NeoRecall does not install or run a transcription or language model. Provider settings can come from .env or from encrypted live overrides on the app's Admin βΊ Providers page. Keys saved there are never returned to the app, and Use the serverβs own configuration restores .env as the source of truth.
For transcription, choose openai, groq, deepgram, assemblyai, or openai-compatible. Set TRANSCRIPTION_API_BASE_URL, TRANSCRIPTION_API_MODEL, and the selected provider's API key. The generic OpenAI-compatible adapter accepts either a version root ending in /v1 or the full /audio/transcriptions URL, sends the audio as multipart field file, and supports optional TRANSCRIPTION_API_LANGUAGE plus TRANSCRIPTION_API_RESPONSE_FORMAT. A model is optional for custom endpoints that route it server-side.
Each account chooses a language under Settings β General, which sets both the language of the app and the language the model writes memories, summaries and answers in. On a fresh installation it is taken from the device's own language and stored; after that only the picker changes it, and the choice follows the account to other devices. English and German are supported. Recordings are still transcribed in whatever language was actually spoken, and search stays multilingual.
Each account can add custom vocabulary under Settings β Recording β Transcription, one word or phrase per line. NeoRecall also includes speaker names the user has explicitly confirmed and shows those names separately in Settings. Names inferred by the language model remain available as display metadata but are deliberately excluded from transcription vocabulary, preventing an incorrect inferred name from biasing later recordings and creating a feedback loop. Existing names are treated as unconfirmed when this provenance tracking is introduced; saving a speaker name in the Speakers screen confirms it. The combined trusted list is sent as an OpenAI-compatible prompt, Deepgram keywords/keyterms, or AssemblyAI keyterms according to the selected provider. For prompt-only OpenAI-compatible providers, an account-level switch controls an additional conservative correction: NeoRecall rewrites a returned word only when it is a long, close, unambiguous match for one single-word vocabulary entry; multi-word phrases are never rewritten. Applied corrections are counted in server logs without logging transcript text or vocabulary. NEORECALL_CUSTOM_VOCABULARY_MAX_TERMS and NEORECALL_CUSTOM_VOCABULARY_MAX_TERM_LENGTH control the list limits. The conservative fallback is configured with NEORECALL_VOCABULARY_CORRECTION_MIN_LENGTH, NEORECALL_VOCABULARY_CORRECTION_MAX_DISTANCE, NEORECALL_VOCABULARY_CORRECTION_SIMILARITY, and NEORECALL_VOCABULARY_CORRECTION_AMBIGUITY_MARGIN.
For generation, choose openai, anthropic, google, groq, mistral, xai, deepseek, openrouter, together, or openai_compatible. Set AI_API_MODEL and either the provider-specific key from .env.example or AI_API_KEY. Custom OpenAI-compatible endpoints also require AI_API_BASE_URL.
Backupsβ
The live database file is page-encrypted with the installation key (data/secret.key). A first start on an older plaintext file encrypts that file in place. NeoRecall takes a scheduled snapshot of its database using SQLite's online backup API, encrypts that snapshot again with the same key, and writes it to the configured destination. Backups are on by default, run every NEORECALL_BACKUP_INTERVAL_HOURS (default 24), and NEORECALL_BACKUP_RETAIN (default 3) artifacts are kept β older ones are pruned automatically. Files in the backup directory that NeoRecall did not write are never touched.
NEORECALL_BACKUP_DESTINATION selects where artifacts land. local (the default) writes to ~/.neorecall/backups. Artifacts are encrypted before they leave the process, so a destination never handles plaintext.
Admin βΊ Backups in the app shows the schedule, the last run, retention, and every past run including failures, and offers a Back up now button.
Each account can also copy its own data to a self-hosted Nextcloud instance under Settings β Integrations. That path is write-only (MKCOL and PUT): it is not a restore source and it is not the admin database backup. Audio copies wait until a recording ends, then upload one joined file rather than each ingest chunk. NEORECALL_CLOUD_USER_BACKUP_INTERVAL_HOURS (default 24) is how often an account with the data-backup toggle on is offered a dump. Pending audio copies older than NEORECALL_CLOUD_PENDING_MAX_AGE_MS are dropped.
From the command line:
neorecall backup # snapshot now
neorecall backup list # what is stored, and the schedule
neorecall restore <key> # decrypt one artifact beside the live database
restore refuses to run while NeoRecall is up, and never writes over the live database. It decrypts the artifact next to the original, runs an integrity check, reports the account count and checksum, and prints the two mv commands to put it into service. Verify a restore periodically β a backup that has never been restored is an assumption, not a control.
Admin βΊ Providers fetches each provider's current model catalog through its API instead of shipping a fixed model list. Providers without a model-list endpoint may route automatically, and custom compatible endpoints remain manually editable if they do not implement GET /models.
LLM_CONTEXT_SIZE describes how much the configured external model can read at once. Consolidation splits longer transcripts to fit it, so raising it buys fewer and wider passes rather than deciding what can be processed at all. The embedding model used for local semantic search remains pinned and verified by neorecall setup; it is not a generation model.
Memory consolidationβ
Processing gates are off by default and remain available when an external deployment cannot keep up or a hosted provider needs tighter request limits.
NEORECALL_MIN_CONSOLIDATION_INTERVAL_MS defaults to 0, so a conversation is consolidated on the scheduler tick after it closes. NEORECALL_MIN_AI_AUDIO_MS and NEORECALL_MIN_NEW_MATERIAL_CHARS default to 0 and 1: a thirty-second exchange is worth describing as soon as it ends. NEORECALL_MAX_CONSOLIDATION_LATENCY_MS defaults to 0, so nothing waits for a batch to fill. Raising any of them restores the old behaviour exactly β NEORECALL_MIN_AI_AUDIO_MS in particular is still a hard floor rather than a heuristic: at one minute, a recording of a minute or less reaches no model at all, not through consolidation, not through a live preview, and not by asking for one by hand.
One occasion, one cardβ
A conversation boundary is a provisional grouping of speech, not an occasion. NEORECALL_CONVERSATION_HARD_GAP_MS cuts the stream after ten minutes of quiet β long enough that a pause reaching it really was the end of the sitting, where the earlier three minutes cut a meeting at every coffee. Shorter pauses only cut where the speech after them is also about something else, so a sitting still arrives as more than one conversation when the subject genuinely moved on. A run therefore carries one occasion: the oldest conversation still waiting, plus every conversation that follows it on the same recording without a longer break.
NEORECALL_MEMORY_OCCASION_GAP_MS defaults to fifteen minutes and is what "without a longer break" means. Below it, two consecutive conversations of one recording are read as one sitting and consolidated together; above it, the later one starts its own occasion. It is never an arbitrary batch: the chain stops at the first conversation from another recording or beyond this gap, so the model is not asked to hold two unrelated occasions in mind at once.
NEORECALL_MEMORY_SETTLE_MS defaults to ten minutes and decides when a sitting is over. While the recording is still running, an occasion is written up only once it has been quiet this long β writing up the first fragment immediately is what produced several cards, minutes apart, for one meeting. Once the recording has stopped nothing waits at all: a stopped recording is proof the occasion ended, and a conversation that just finished is the one you are about to look for. Asking by hand also skips the wait. Both this and the occasion gap must be at least NEORECALL_CONVERSATION_HARD_GAP_MS, as must NEORECALL_CONVERSATION_QUIET_CLOSE_MS; the server refuses to start otherwise.
NEORECALL_MEMORY_OCCASION_MAX_WAIT_MS defaults to one hour and bounds the wait. An always-on recording never stops, so at this age an occasion is written up with what it has, and later fragments reach the continuation mechanism below.
NEORECALL_MAX_CONSOLIDATION_CONVERSATIONS defaults to 12 and is the ceiling on one chain, not a batch size. What actually bounds a request is NEORECALL_MAX_CONSOLIDATION_INPUT_CHARS and NEORECALL_CONVERSATION_MAXIMUM_MS; a chain cut by any of them is finished by the continuation mechanism instead. After a validation failure the next run carries a single conversation regardless, so the cause can be attributed.
NEORECALL_MAX_MEMORY_CONTINUATION_CANDIDATES defaults to 8. A run still reads one new provisional conversation, but it also shows the model a bounded set of recent or same-recording memory cards. The model must explicitly identify which, if any, are fragments of the same real-world occasion. Claimed fragments are updated or absorbed into one card while their transcript sources and existing highlights remain attached. Time and recording continuity only narrow the candidates; matching titles, keywords, or a similarity threshold never decide a merge, and a recurring lesson or meeting remains separate unless the model identifies it as the same continuous occasion.
NEORECALL_MEMORY_CONTINUATION_LOOKBACK_MS defaults to two hours. It controls how far back cards from a different recording session remain available for that decision, covering recorder restarts, reconnects, and delayed sync without treating a three-minute conversation boundary as a merge horizon. Same-stream cards remain eligible independently of this window. The prompt also receives recurring-speaker overlap when it is available; neither overlap nor time performs a merge on its own.
Merging duplicates that got throughβ
Consolidating a whole occasion at once handles the ordinary case, but it can only join what it can see. A device that reconnects starts a new recording stream, an occasion longer than the maximum wait is written up before it ends, and a fragment whose transcription finished late arrives after its neighbours were already written. Each leaves two cards for one sitting, so a sweep in the maintenance job looks for them.
NEORECALL_MEMORY_DEDUPE_ENABLED defaults to true. Each pass reads the cards written since the last pass, finds those close enough in time and alike enough in wording to be worth a question, and asks the model whether they describe the same occasion. Cards it says are the same are folded into one, keeping every highlight, topic, entity and transcript line; a card whose wording you edited yourself keeps your words and is not rewritten.
NEORECALL_MEMORY_DEDUPE_WINDOW_MS defaults to six hours and is the guard that matters most: two lessons of one course or two calls about one project read almost identically, and only time separates them, so nothing outside this window is ever considered. NEORECALL_MEMORY_DEDUPE_SIMILARITY_THRESHOLD defaults to 0.88 and decides which pairs inside it are worth asking about β nothing is merged on this number alone. Multilingual-e5 similarities sit high even between unrelated text, so set it against your own recordings rather than by intuition: node scripts/memory_dedupe_report.js prints the pairs and where their scores fall without asking the model anything or changing a card. NEORECALL_MEMORY_DEDUPE_MAX_PAIRS_PER_RUN (default 20) and NEORECALL_MEMORY_DEDUPE_NEIGHBOURS (default 5) bound what one sweep can cost.
NEORECALL_MEMORY_MERGE_MAX_ITEMS defaults to 100 and bounds a manual merge request. The server advertises the effective value to clients, combines the selected evidence immediately, and leaves the optional title and summary rewrite to a background job.
NEORECALL_MIN_MEMORY_EVIDENCE_MS and NEORECALL_MIN_MEMORY_EVIDENCE_CHARS are unchanged and are what keeps short speech off the timeline as a memory card (defaults: two minutes of speech and 400 transcript characters). Below either floor the section still receives a title and summary, but it is not memory-worthy. Mini-memories under a larger worthy occasion are reserved for concrete, still-open action items with an identifiable owner; facts, observations, suggestions and unaccepted requests stay in the memory summary instead. The consolidation prompt states the same bar; the floors enforce it when the model over-promotes short speech.
A consolidation retries only failures that say nothing about its input β no message content, a timeout, a transport error β bounded by AI_MAX_RETRIES. An answer that violates the contract is never resent unchanged, because resending reproduces it; narrowing and quarantine handle that case instead. Ask uses its own NEORECALL_ASK_MAX_PER_HOUR database quota and minute burst limiter so one client cannot overwhelm the configured provider while recordings are still arriving.
Rolling per-user provider budgets sit beside those Ask counters. They are off by default (0 = unlimited) so a self-hosted install does not suddenly stop processing:
NEORECALL_AI_TOKENS_4H/NEORECALL_AI_TOKENS_WEEKLYβ language-model tokens over a rolling 4-hour and 7-day windowNEORECALL_TRANSCRIPTION_SECONDS_4H/NEORECALL_TRANSCRIPTION_SECONDS_WEEKLYβ audio seconds that actually went to the transcription service (local silence detection does not count)
Admin βΊ Users in the app can set the same four install defaults without writing .env, and a Limits control on each account can inherit them, replace them, or set 0 for unlimited. When a cap is reached, Ask returns 429 USAGE_LIMIT_EXCEEDED, memory writing and previews wait, and transcription of speech is deferred. The job is not failed: attempts are not burned, the server keeps its temporary audio, and no terminal receipt is issued, so the recording device keeps the original until the window opens.
The day's summaryβ
The daily summary is written by its own small request once the transcript has been read, not by the pass that reads it. With windowing no single pass sees the day, so asking one to summarise it asks it to write about material it was never shown. The separate request reads only the titles and summaries of the memory-worthy sections the run produced.
Which local date that summary covers, and in which timezone, are derived from the conversations the run selected rather than read back from the model. The server already knows both, so asking the model to restate them only created a way for the answer to disagree with the evidence β and that disagreement used to fail an otherwise correct consolidation.
Windowing a long transcriptβ
A four-hour lecture does not fit in a typical model context, and it must still become one memory. Consolidation therefore splits the transcript into windows that each fit LLM_CONTEXT_SIZE, cut on segment boundaries and processed in order. Each window after the first is told what the occasion looked like when the previous window stopped and marks the section β and the memory built from it β that carries on, so the two are folded back into one. A transcript that fits is exactly one request and behaves as it always did.
NEORECALL_CONSOLIDATION_WINDOW_CHARACTERS is how much transcript one window carries, and it is sized against the answer rather than against the context. Those are different quantities, and the answer is the one that fails: a full contract for dense speech runs to roughly one output token per five input characters, so a window sized to fill a 16 384-token context β nearly thirty thousand characters β asks for several times more answer than AI_CONSOLIDATION_MAX_OUTPUT_TOKENS allows and arrives truncated. The default of 8 000 characters is five to eight minutes of speech and leaves the answer a fourfold margin. Raising it lets the model see more of an occasion at once; lowering it is the first thing to try if AI_OUTPUT_TRUNCATED appears. It is clamped to whatever the context can hold, so it can never exceed LLM_CONTEXT_SIZE minus the output budget.
AI_CONSOLIDATION_MAX_OUTPUT_TOKENS bounds the answer for one window and shares the context budget with the prompt, so it cannot be raised without raising LLM_CONTEXT_SIZE too. NEORECALL_MAX_CONSOLIDATION_INPUT_CHARS still bounds what one run may carry before windowing splits it.
When the context runs out anywayβ
LLM_CONTEXT_SIZE is a claim about somebody else's server, so it can be wrong. Set it larger than the endpoint really allows and every request overflows β and an overflow is not a transport fault: it produces the identical rejection however many times it is sent. Treated as transient it would be retried, fail the run without narrowing or quarantining anything, re-enter the candidate set on the next scheduler tick, and repeat indefinitely without ever producing a memory.
NeoRecall therefore reads the rejection rather than only its status code. Every vendor words it differently β context_length_exceeded, maximum context length isβ¦, prompt is too long, exceeds the available context β and any of them becomes AI_CONTEXT_EXCEEDED, which is sent once, never retried, and narrows the batch exactly like any other input the model could not handle. An ordinary bad request (a rejected key, a rate limit) is untouched and still retried.
The error names the two settings that fix it. Lower LLM_CONTEXT_SIZE to what the endpoint actually allows; that alone re-sizes every window. If the endpoint is small enough that the output budget no longer leaves room for a prompt, the server refuses to start rather than sending requests that cannot fit, and AI_CONSOLIDATION_MAX_OUTPUT_TOKENS has to come down with it.
Every other path that builds a prompt is bounded the same way. Ask trims retrieved evidence to fit, dropping the weakest matches first, because search returns its results best-first and answering from slightly less evidence beats being refused. The day's summary trims the occasions it reads, which matters because a long recording produces many sections. Live previews already cut to a budget, falling back to the next refresh for whatever did not fit.
What one request may returnβ
The contract caps a single pass at three memories, eight mini-memories per memory and sixteen entities. Without those bounds a model handed a dense transcript can emit one mini-memory per utterance until it exhausts its token budget.
The caps bound a window, not an occasion. A three-hour lecture is read in many windows whose results merge, so it still accumulates as many mini-memories as it deserves while no single request grows without limit. Conversation sections are deliberately uncapped: they have to partition the whole input.
Because the caps make the largest possible answer arithmetic rather than a guess β three memories with eight mini-memories each, sixteen entities and the sections around them come to roughly five and a half thousand tokens β AI_CONSOLIDATION_MAX_OUTPUT_TOKENS can be sized to cover it. Its default of 8 000 does, with margin; measured runs of a dense 8 000-character window landed between 2 400 and 3 900.
Throughputβ
Provider latency depends on the selected deployment and model. AI_REQUEST_TIMEOUT_MS defaults to thirty minutes and NEORECALL_JOB_LEASE_MS matches it: the timeout has to outlast the slowest legitimate answer, and a lease shorter than the job would let a second worker start the same run while the first is still writing.
All of it is background work, but preview intervals and per-run conversation limits still determine provider load.
When an answer does not fit the contractβ
NeoRecall requests structured JSON and validates every response before changing memory state. Missing fields, invented enum values, invalid references, and prose around the JSON fail validation. A completion that reaches its provider token limit is reported as AI_OUTPUT_TRUNCATED; both truncation and contract failures narrow the next run rather than silently dropping evidence.
Candidates are built oldest-first, so a conversation the model cannot partition would otherwise reappear in every later run. After a validation failure the next run carries a single conversation, and NEORECALL_CONSOLIDATION_MAX_FAILURES bounds how often one conversation may fail before it is quarantined. A quarantined conversation keeps its transcript and stays readable but no longer blocks memory generation.
Live conversation previewsβ
A conversation that is still being recorded gets a provisional title, summary and topics so it can be read before it ends. AI_PREVIEW_MAX_OUTPUT_TOKENS bounds the completion; a preview answer is three short fields, and the default model does not spend tokens thinking first.
Preview work is bounded by transcript growth rather than by elapsed time: NEORECALL_CONVERSATION_PREVIEW_MIN_CHARACTERS is how much transcript the first preview needs, NEORECALL_CONVERSATION_PREVIEW_REFRESH_CHARACTERS how much new transcript each refresh needs, and NEORECALL_CONVERSATION_PREVIEW_MIN_INTERVAL_MS the minimum spacing between two previews of the same conversation. They sit close to the scheduler tick β 300 characters, 600 characters, one minute. Raise them if the provider falls behind. The interval is measured from the last attempt rather than the last success, so a model that cannot satisfy the contract costs one request per interval instead of one per scheduler tick.
Beyond NEORECALL_CONVERSATION_PREVIEW_FULL_CHARACTERS a refresh sends the previous description plus only the speech recorded since, so a conversation that runs all day takes the same work per refresh instead of re-reading its whole history. A request is finally cut to what the model can read at once; when that bites, the description continues on the next refresh instead of the request failing. Any drift those rolling summaries accumulate is corrected when the conversation closes and consolidation reads the full transcript.
Previews never create memories; consolidation replaces the insight and marks it final when the conversation closes. A quarantined conversation is the exception: consolidation will never describe it, so previews keep it readable instead of leaving an unlabelled transcript in the timeline.
NEORECALL_SCHEDULER_INTERVAL_MS is how often the worker looks for work, and therefore the coarsest term in how long after crossing a threshold a result appears.
NEORECALL_IMPORT_SESSION_CONTINUITY_MS is how large a gap may be between two imports from one device before they stop counting as the same recording stream. It has to comfortably exceed the client's device-sync poll and its failure backoff.
Audio conditioningβ
Recordings reach the server from pocket wearables with millimetre microphones, from meeting bots, from Discord and from files somebody imported, and they arrive tens of decibels apart with whatever rumble and hiss the room contributed. Before anything listens to a chunk, a short ffmpeg filter chain levels and cleans it. It is ordinary signal processing β a high-pass, gentle spectral denoising, level normalization, a limiter β and none of it knows which language is being spoken, so it helps every language the same way.
Nothing in the chain changes how long the recording is. That is not a preference: speaker turns, transcript timestamps and the voice previews cut later from the original chunk all describe one timeline, and a stage that added or removed audio would slide them apart silently. Nothing that trims, gates or stretches belongs here.
If conditioning fails for any reason β an unreadable chunk, a missing filter, a deadline β the original recording is transcribed instead and a warning is logged. The worst this feature can do is nothing.
NEORECALL_AUDIO_PREPROCESS_ENABLED=false turns it off entirely.
NEORECALL_AUDIO_PREPROCESS_HIGHPASS_HZ (70 Hz) removes rumble, handling noise
and any DC offset the capture device introduced; no language carries meaning
that low. 0 disables the stage.
NEORECALL_AUDIO_PREPROCESS_DENOISE_DB (6 dB) is deliberately gentle. Strong
noise reduction smooths the onset of plosives and fricatives, which is exactly
the detail an acoustic model reads, so it can cost more accuracy than the noise
did. Raise it only for consistently noisy recordings, and compare the
transcripts afterwards. NEORECALL_AUDIO_PREPROCESS_DENOISE_ENABLED=false
removes the stage.
NEORECALL_AUDIO_PREPROCESS_NORMALIZER decides how the level is evened out.
dynaudnorm is the default and costs almost nothing. loudnorm is the
broadcast-correct answer and roughly ten times more expensive, because it
resamples internally to measure true peaks; it is also the only mode that
reports the loudness it measured, which the log then carries.
speechnorm is more aggressive and lifts the noise between words along with the
speech. off leaves levels alone. NEORECALL_AUDIO_PREPROCESS_MAX_GAIN caps
how far a quiet passage may be lifted, so a near-silent room's noise floor is
never amplified into something that looks like speech.
NEORECALL_AUDIO_PREPROCESS_FORMAT=flac roughly halves what is uploaded, at no
loss, for a metered connection β conditioned audio is uncompressed by default,
which is several times the size of the Opus a wearable sends.
NEORECALL_AUDIO_PREPROCESS_MAX_DURATION_MS sends unusually long recordings
straight to the service instead of holding them in ffmpeg.
audioPreprocessEnabled, audioPreprocessHighpassHz, audioPreprocessDenoiseDb
and audioPreprocessMaxGain are also processing settings, so they can be
changed on a running server without a restart.
Speaker detection reads the original recordingβ
NEORECALL_AUDIO_PREPROCESS_TARGET decides which passes hear the conditioned
audio. The default, stt, gives it only to the transcription service; speech
detection and speaker identity keep reading the recording as it arrived.
That default is a measurement, not caution. The segmentation and speaker-embedding models were trained on unprocessed speech, and on the two-speaker test fixture every conditioned variant separated the voices worse than the raw audio did. The high-pass on its own was the most damaging: it merged both people into a single speaker, because part of what distinguishes one voice from another lives in exactly the low frequencies it removes. Conditioning helps a transcription service and hurts these models, so each gets the audio it does better with β which is safe only because conditioning does not move the timeline, so both passes still describe the same recording.
stt+analysis gives the conditioned audio to all of them. There is a second
cost to it beyond the above: a voice fingerprint is only comparable to one taken
under the same conditions, so people enrolled before the switch can start reading
as somebody new. The server says so once at start-up when it finds enrolled
voices, and the Speakers screen's re-detect re-resolves recent recordings under
the current conditions and folds duplicate profiles back together.
If you do try it, compare the speaker labels on a recording you know before leaving it on.
Speech detection and speaker identityβ
These two run on the audio itself, in the NeoRecall process, and they are the only inference it still does. The reason is not size but capability: a transcription service returns what was said, never who said it. The best it offers is a speaker label valid inside the single request that produced it, and since every chunk is its own request, that label cannot be carried across a chunk boundary. A voice embedding can β it is what makes a speaker the same person in a recording made next week.
They are cheap enough for that to be uncontroversial: a 640 KB voice-activity
detector and 31 MB of segmentation and speaker-embedding weights, against the
gigabytes recognition and generation would need. neorecall setup installs them
with everything else.
Speech detection earns its place twice. A chunk it finds no speech in is never
sent to the transcription service at all, so an idle microphone costs nothing β
which on a recorder running all day is most of the day. NEORECALL_VAD_THRESHOLD
sets that bar; raise it to send less, lower it if quiet speech is being missed.
NEORECALL_VAD_MIN_SPEECH_SECONDS and NEORECALL_VAD_MIN_SILENCE_SECONDS decide
how readily it opens and closes a span.
NEORECALL_DIARIZATION_ENABLED=false turns both off. Every chunk then goes to
the service, silence included, and transcripts arrive with no speaker labels.
NEORECALL_SHERPA_THREADS bounds the native threads; the models run per chunk on
a handful of seconds of audio, so the default of two is deliberate.
Testing it end to endβ
Admin βΊ Providers has a Save and test button. It sends a bundled eighteen-second sample of real two-speaker speech to the configured transcription service, runs the local speech-detection and diarization pass over the same sample, and makes one structured request to the language model β then reports the three legs separately, with the words that came back.
Separately on purpose: a transcript that never arrives, a transcript with no
speakers, and a model that refuses the JSON contract need three different fixes,
and a single "failed" would hide which. The language-model leg uses the same
budget a live preview gets, so it measures the pipeline rather than a stricter
version of it. Transport failures name their cause rather than reporting
fetch failed, so a wrong port reads as a wrong port.
Models that think before answeringβ
A reasoning model bills its internal deliberation against the same output budget
as its answer, and spends it first. A request for a three-line preview can
therefore exhaust the budget mid-thought and return nothing usable β which arrives
looking exactly like a model that cannot follow the contract, and sends you to
rewrite prompts instead of raising a number. AI_OUTPUT_TRUNCATED now reports the
reasoning share when the endpoint provides it, so the cause is visible.
Two remedies. Raise AI_PREVIEW_MAX_OUTPUT_TOKENS and
AI_CONSOLIDATION_MAX_OUTPUT_TOKENS β headroom is free, since max_tokens is a
ceiling rather than a target and a model that answers briefly still stops briefly.
Or switch the thinking off, usually the better trade on a small local model where
deliberating costs minutes per request. That field is not standardised, so it goes
through AI_API_EXTRA_BODY, a JSON object merged into every request:
AI_API_EXTRA_BODY={"chat_template_kwargs":{"enable_thinking":false}}
Qwen-family servers accept that spelling; others differ, which is why this is a passthrough rather than a setting per vendor. It merges last, so it can override anything NeoRecall sets.
Admin βΊ Providers carries the same thing without the typing: under the language model's Advanced, a Skip the modelβs thinking step switch that writes exactly that payload into the Extra request JSON field beside it. The switch is a view of one key rather than a separate setting β turning it on alongside JSON you wrote yourself adds the key, and turning it off removes only that key. Saved there it is stored with the rest of the provider settings and takes effect without a restart.
You should rarely need it, because a truncated completion rescues itself: when a
request runs out of budget mid-thought, NeoRecall retries it once with the
thinking step off. That is deliberately narrow. It happens only after a real
truncation, only once, and only when you have not set the field yourself, so a
provider that has never heard of chat_template_kwargs is never sent it
speculatively β and if the retry is itself rejected, the original truncation is
what gets reported, since that is the fault worth fixing.
The rescue matters most where truncation hurts most. Consolidation treats a truncated answer as the input's fault: it narrows the batch, and after enough failures quarantines the conversation. Without the retry, a model that always deliberates would work through an entire backlog that way, quarantining recordings that were never the problem.
How a transcript gets a speakerβ
The local pass and the service work on the same seconds of audio and meet on the same timeline. Diarization produces speaker turns with an embedding each; the service returns timestamped text for the whole chunk; every segment takes the turn it overlaps most, and that turn's embedding with it. A segment with a second voice talking across more than a fifth of it is marked as overlapping rather than quietly credited to one person.
Before persistence, a phrase-independent quality guard also catches the failure
mode where an ASR provider fills a short timestamp with the same word or short
token template dozens of times. It compacts a run only when it repeats at least
NEORECALL_TRANSCRIPT_REPETITION_MIN_REPEATS times (default 8), covers at least
NEORECALL_TRANSCRIPT_REPETITION_MIN_COVERAGE of the segment (default 0.8),
and would require more than NEORECALL_TRANSCRIPT_MAX_WORDS_PER_SECOND (default
5) to have actually been spoken. Patterns are bounded by
NEORECALL_TRANSCRIPT_REPETITION_MAX_PATTERN_WORDS (default 8). Numeric slots
may vary between repetitions, which handles counter-like hallucinations without
keying the detector to a word such as a speaker label. One occurrence remains in
the transcript as evidence; normal emphasis, lists, and realistically paced
repetition are preserved.
A voice is fingerprinted from everything it said in the chunk, weighted by how
long each turn lasted β not from a single turn. That matters more than any
threshold. Measured against two known-different voices: with one second of speech
per fingerprint the same voice scored anywhere from 0.20 to 0.87 while two
different voices reached 0.77, so the two populations are indistinguishable. With
two seconds they separate cleanly β the same voice never below 0.55, different
voices never above 0.50. NEORECALL_SPEAKER_CLUSTER_THRESHOLD sits at 0.52,
inside that gap. NEORECALL_SPEAKER_MINIMUM_TURN_MS defaults to 500: this
deliberately favors giving real short speech a possibly imperfect speaker label
over leaving it unlabeled, while still preventing the briefest diarization blips
from founding a profile. Below that duration a turn may join a voice that already
exists, but it cannot invent one.
A match needs NEORECALL_SPEAKER_CLUSTER_MARGIN over the runner-up only when that
runner-up is itself below the threshold β the case the margin exists for, where a
new speaker grazes the bar against an unrelated cluster and both readings are
equally weak. When two clusters both match strongly they are not an ambiguity to
refuse but one person split earlier, so the best match wins and the two are merged
if their centroids are at least NEORECALL_SPEAKER_CLUSTER_MERGE_THRESHOLD alike.
Refusing both used to mint a third copy, which made the next turn more ambiguous
still β a loop that could turn one familiar voice into a dozen entries.
Diarization restarts on every chunk, so a speaker crossing a boundary can drift
below the plain threshold with nothing about the voice having changed: when a
component's first speech begins within NEORECALL_SPEAKER_CONTINUITY_GAP_MS of
where its last known turn ended, that cluster may be kept at the relaxed
NEORECALL_SPEAKER_CLUSTER_CONTINUITY_THRESHOLD. That only ever breaks a near-tie,
so a genuine speaker change at the boundary still resolves on its own.
Enrolling a durable voice β a person recognised across recordings, rather than a
cluster inside one β is the only speaker decision more evidence cannot undo. A
spurious profile is permanent, appears as its own unnamed person, and then
competes for every later match, so it takes more than merely failing to match a
known one. A voice must score below NEORECALL_VOICE_ENROLL_FLOOR (default
0.45) to count as somebody new, and its fingerprint must be pooled from at
least NEORECALL_VOICE_ENROLL_MIN_MS of speech (default 3000) β below that, a
low score says the measurement was poor, not that the voice is unknown. Between
the floor and NEORECALL_VOICE_MATCH_THRESHOLD lies a grey band where a voice
resembles someone enrolled without confirming it; there, and where a match over
the bar is within NEORECALL_VOICE_MATCH_MARGIN of a candidate below it, the
turn is attributed to nobody and left for a later chunk with better evidence. It
still carries its conversation-local speaker label throughout. Two profiles that
both clear the bar are one person already split rather than an unclear reading,
so the stronger wins β the same correction the cluster layer describes above, for
the same reason: refusing both used to mint a third copy, which made the next
turn more ambiguous still.
Speaker embeddings are not unit vectors, and their magnitude tracks loudness and turn length rather than who was talking, so every sample folded into a profile is normalized to a direction first. Left raw, one loud or long sample drags a profile off the voice it stands for until the person stops matching themselves and a duplicate is minted. The weight of a profile's accumulated history is capped, so a voice first enrolled through one microphone can still migrate toward the same person heard through another instead of freezing around whatever the first conversation sounded like.
Recurring matching also reconciles duplicate profiles automatically after new
speech is persisted and during hourly maintenance, so profiles already present
when a server is upgraded are cleaned up as well. Mutually nearest profiles above the configured voice-match
threshold are folded together unless their explicit names or linked person
entities conflict. The Speakers screen's re-detect goes further, because there
the user has looked at the list and asked for the duplicates to be sorted out: it
also folds together a pair below the match threshold but above
NEORECALL_VOICE_REPAIR_THRESHOLD (default 0.50) when each profile is the
other's closest match and stands clear of its own runner-up by the voice-match
margin. Mutual exclusivity is far stronger evidence than a one-sided score; mere
adjacency in a crowd of profiles is not, and is refused. The automatic pass after
each chunk stays at the strict bar. Inside one recording, session clusters that resolve to the
same recurring voiceprint are collapsed immediately; while that derived cleanup
runs, all such clusters already share one conversation-local label. Cluster
cleanup is best-effort and never delays transcript persistence, server-side
audio deletion, or the terminal receipt.
If one person still appears as several, lower NEORECALL_SPEAKER_CLUSTER_THRESHOLD;
if different people are being merged, raise it. Note that
NEORECALL_DIARIZATION_CLUSTER_DISTANCE, which groups voices inside a chunk,
points the other way β it is a distance, so raising it yields fewer speakers. They
were one setting until they were found to be pulling in opposite directions.
When the recording already knows who is speakingβ
Everything above is inference, and inference is what produces the same person twice. Some recordings never needed it. A capture path that receives a separate stream per participant is told whose stream it is, and working that out again from the sound of the voice is both slower and worse than the fact it was handed.
Any source can say so, in the metadata it already sends when it registers:
{ "speaker": { "key": "chat:1234", "name": "Mara" } }
key is opaque and scoped to the user β namespace it ("<source>:<id>") so two
sources cannot collide. Nothing in the server knows what produced a key, and no
integration is named anywhere in that path: a source that knows, says so, and one
that does not says nothing and is matched acoustically exactly as before. This is
why it works the same for a chat bot, a per-participant recorder, or a client
that simply knows it is one person's headset.
A declared stream skips voice matching entirely. The person is resolved first and
the recording-local voice follows from them, which prevents two failures rather
than repairing them: one person cannot split into several labels however badly a
fingerprint was measured, and two people who happen to sound alike cannot collapse
into one. name only ever fills a gap β it is stored as inferred, so a name the
user sets, or one consolidation reads out of a self-introduction, always wins.
The second benefit is larger than the first. A declared stream is labelled speech, which acoustic matching never gets, so the profile it builds is correct by construction β and that profile is then what recognises the same person on a room microphone or a pendant, where nothing is declared. Fingerprinting is withheld for any chunk that turned out to carry more than one voice: the stream still belongs to the person it names, so their label stands, but a second person audible behind an open microphone must not be learned as them.
npm run speakers:report shows how many profiles were identified this way rather
than by voice. A high share is the cheapest accuracy an installation can have.
Looking again once the conversation endsβ
Everything above decides who is speaking from one chunk of audio at a time,
because while a recording is running that is all there is. Two costs follow. A
voice re-segmented at a chunk boundary can drift below the matching bar and start
a second identity; and a fingerprint pooled from a few seconds is often too
little speech to enroll anyone at all, so its turns attach to no durable person β
and with nothing to group them by, each cluster keeps its own Speaker N. That
is the "one person, three labels" complaint, and neither cost is a threshold that
could be tuned away. Both are consequences of having to decide early.
When a conversation closes, that constraint is gone: every voice in it is on the
table, and the speech behind each one is conversation-scale rather than
chunk-scale, which clears the enrollment floor per-chunk speech usually cannot.
Closing therefore queues one pass that re-asks the question with the whole
conversation in view. It groups the conversation's voices by average-linkage
similarity at NEORECALL_SPEAKER_CLUSTER_MERGE_THRESHOLD, refusing any pair
heard speaking over each other on the same recording β the one thing that proves
two voices are two people β and any pair already carrying different names the
user set. Then it resolves each group to a person once, with all of its speech
behind the decision. Turned off per user with Review speakers when a
conversation ends; NEORECALL_SPEAKER_REDETECT_DAYS (default 30) bounds how
far back the Speakers screen's re-detect re-runs it.
Labels can therefore change shortly after a recording finishes. That is safe
because of when it happens: consolidation only ever reads conversations that have
closed, and is held back from any conversation still waiting on this pass, so
corrections land before the model reads a speaker label rather than contradicting
something already written. Clusters are folded together only within one
recording session β speaker_clusters is unique on its session and the live
resolver looks a cluster up by session, so a row moved out of its own session
would become invisible to the recording still producing it, which would mint a
replacement every chunk. The same voice heard in two sessions is given the same
voiceprint instead, which collapses the label just as well and destroys nothing.
Reconciling duplicate profiles is queued rather than run inline for the same reason it exists at all: it compares every enrolled voice against every other one inside a transaction, and doing that once per chunk put the cost on the path that has audio waiting to be deleted and a receipt waiting to be issued.
A conversation queues its own resolution when it closes, so on a healthy server nothing else is needed. Because nothing otherwise ever revisits a closed conversation, hourly maintenance also sweeps for conversations that finished but still contain speech belonging to no durable person β the state this pass exists to correct β and queues those. That covers a worker that was down at the wrong moment, a job that ran out of attempts, and conversations recorded before any of this existed. Both the sweep and re-detect are capped per run so an upgrade turns into a steady backlog rather than a stall; whatever is not queued this hour is queued the next.
Speech belonging to nobody is not on its own proof that the pass has not run. Speech far too short to found a person enrolls nobody, and a voice that resembles an enrolled one without clearing the bar is left alone rather than guessed at β both are finished answers that leave the speech unattached. The pass therefore records what it concluded about each conversation, and the sweep skips a conversation while that answer still applies; without it the sweep would queue the same unresolvable conversation on every maintenance tick forever, each job completing without changing anything. The answer stops applying, and the conversation is looked at again, as soon as something that could change it has: a voice enrolled, deleted, merged, renamed or re-enabled, one of the thresholds above moved, a new version of the pass shipped, or more speech landing in the conversation itself.
npm run speakers:report prints what this actually looks like on an
installation β how many profiles are duplicates of each other, how much speech
carries no durable identity, and where the similarity scores really fall.
Thresholds here were chosen against a two-speaker measurement, which is a guess
about anybody else's recordings; run the report before changing one and again
afterwards.
Consolidation then identifies people from the transcript β a self-introduction, or another speaker naming them β and names that speaker's voiceprint from the same response, at no extra request. It never overwrites a name set by hand.
When it is unavailableβ
The native runtime ships prebuilt binaries for the common platforms and is an
optional dependency, so an install on a platform it does not cover has no audio
models. That is survivable and deliberately not fatal: chunks are transcribed
exactly as before and segments simply carry no speaker. The server reports this
as speakerIdentityAvailable, /ready still passes, and the clients show the
speaker settings switched off with the reason rather than offering choices that
would change nothing.
Operational thresholdsβ
Admin βΊ Processing in the app can safely tune boundary, deduplication, speaker matching, and consolidation material thresholds; only changed values are saved as overrides. Values are validated and stored in SQLite. Environment defaults remain the source of truth until an administrator explicitly overrides a value.
Conversation boundaries expose separate controls for hard and soft silence gaps, contextual embedding similarity and valley prominence, the number of neighboring segments used as semantic context, and maximum duration/character safety ceilings. NEORECALL_CONVERSATION_MAXIMUM_CHARACTERS must not exceed NEORECALL_MAX_CONSOLIDATION_INPUT_CHARS, ensuring one provisional conversation always fits in a bounded consolidation request.
The character ceiling is deliberately set above what NEORECALL_CONVERSATION_MAXIMUM_MS can produce, so duration rather than transcript length is what ends a conversation. A lower ceiling would split a long lecture or meeting into several conversations and therefore several memories purely because it ran long. Lowering it to suit a small-context model is supported, but AI_CONSOLIDATION_MAX_OUTPUT_TOKENS must stay large enough for the sections and memories a full-size input justifies β a completion cut off mid-JSON is indistinguishable from a validation failure.
NEORECALL_SPEAKER_PREVIEW_MIN_MS and
NEORECALL_SPEAKER_PREVIEW_MAX_MS bound the derived clean-speaker sample.
Both are validated within the product's 1β10 second storage contract.
NEORECALL_SPEAKER_DISPLAY_MIN_PREVIEW_MS controls when a profile is mature
enough to appear on the Speakers screen and defaults to a full 10 seconds.
NeoAgent and MCP connectionsβ
NeoAgent connects through NeoRecall's companion OAuth flow. In NeoAgent, open Integrations, select NeoRecall, and enter this server's base URL. The browser then returns to NeoRecall for sign-in and explicit consent.
Claude, ChatGPT (MCP), Cursor, and other MCP clients connect to the same
read-only tools through Streamable HTTP MCP. In NeoRecall, open Settings β
Integrations and copy the MCP URL ({origin}/mcp). Paste that URL into the
client. It registers itself, then the browser returns to NeoRecall for sign-in
and the same explicit consent.
The issued access is limited to search:read, memories:read, and
recordings:read. PKCE is mandatory, refresh tokens rotate on every use, and
the authorization page shows the exact callback URL. Connected apps cannot
upload audio, change memories, start consolidation, or call NeoRecall Ask.
Companion bootstrap and MCP discovery advertise authorize/token/MCP endpoints
using the host the client used to reach NeoRecall, so a LAN or reverse-proxy
hostname works even when NEORECALL_PUBLIC_URL is still localhost. Set
NEORECALL_PUBLIC_URL to the externally reachable HTTPS origin when NeoRecall
is behind a reverse proxy for other public-facing links. Local HTTP URLs remain
suitable when both services run on a trusted private host.
NeoAgent's PUBLIC_URL must resolve in the browser that completes OAuth; that
callback path is always /api/integrations/oauth/callback. MCP clients register
their own HTTPS redirect URIs through /oauth/register.