Scailout Station
Blog
EN
Sign up

Blog

Nine questions a vendor would rather you did not ask

Robert Kim ·

The checklist our own engineers carry into a manufacturer meeting, published in the form we use it.

A demo from a lead capture hardware vendor is designed to be watched, not interrogated. Someone puts a unit on a table, speaks into it, and a transcript appears on a laptop. Everything in that thirty seconds works. Nothing in it answers the questions that matter on day three of a show, with two conversations running at one counter and a card that is 94 percent full.

The questions below are the ones we take to a manufacturer. We wrote them for our own build, they are the list our hardware partner gets before we agree to anything, and there is no reason a buyer should not have the same list.

What to establish before the demo

Establish nine things before you watch anything: record framerate, resolution, audio processing location, microphone clock, indicator control, transfer path, power-loss behaviour, calibration ownership, and retention. Each has a checkable answer. Get them before the demo starts, because a demo answers none of them and is not built to.

A lead capture hardware vendor questionnaire is a set of specification questions a manufacturer must be able to answer in writing, and it is not the same thing as a feature comparison. The distinction matters because features are demonstrable in a controlled room and specifications are not. Every question here shares one property — a vague answer is itself the answer. A vendor who has measured three cameras encoding at once will tell you the number they measured. A vendor who has read a marketing datasheet will tell you the chip is capable of it. Those are different sentences, they cost different amounts of engineering, and they predict opposite outcomes at 4 p.m. on the second day of a show. Take the nine below in the order they are written. Ask for each answer in a document you keep after the meeting ends.

Question Why it matters What a good answer sounds like
Record framerate and inference framerate, separately You can raise inference later, you can never un-record Records 30fps, samples 2fps for analysis
Resolution, and the distance at which a face is identifiable Sets what the device can and cannot recognise A pixel budget, not a resolution number
Where audio is processed Determines who holds the raw signal On device, or a named region and processor
Microphone clock architecture Beamforming is a clock problem first One clock, common I2S/TDM bitclock, synchronous ADC
Whether the recording indicator can be disabled The one question with a legal consequence Blink and minimum brightness are in firmware
Transfer path and measured throughput A day of video is not a small file An interface, a rate, and a duration
Behaviour when power drops mid-recording Shows floors lose power Segment-level durability, named segment length
Who owns calibration data Without it, footage is not reprocessable Per-camera intrinsics shipped with the unit
Retention and deletion policy Determines your own compliance exposure A default, a maximum, and a deletion path
Macro view of the microphone grille a capture device buyer should ask about before signing anything — lead capture hardware v
Macro view of the microphone grille a capture device buyer should ask about before signing anything — lead capture hardware v

Questions about what it records

Record framerate and inference framerate are two separate settings and a vendor who conflates them has not built for this. Recording at a low framerate to save compute is irreversible, because a human reviewing footage later sees stutter that no processing restores. Inference sampling can be raised at any point after the show. Ask for both numbers.

The arithmetic here is easy to re-derive and worth doing in front of the vendor. An eight-hour exhibition day across three cameras is 2,592,000 frames at 30fps. Sampling for analysis at 2fps leaves 172,800 frames, and gating on re-entry detection cuts that again. This is a requirement we calculated to size a batch job, not a measurement of any device: the point is that the analysis load is a scheduling choice while the recording framerate is a permanent one.

Resolution deserves the same treatment. A face-recognition embedding model of the common kind wants roughly 100 pixels across a face, and detection alone survives at 20 to 30. Take an adult face at 16 cm, a 1080p sensor at a 120-degree field of view, and divide: a face at 1.5 m lands near 98 pixels, and one at 3 m near 49. That calculation ignores lens distortion and edge-of-frame falloff, so a real wide lens does worse and the honest version needs field measurement. It is also the reason the distance limit is the privacy feature rather than a specification to apologise for. A vendor proposing 720p for identity work has not run the division.

Questions about where the audio goes

Ask where audio is processed and who holds the raw multichannel signal, because that answer determines your exposure and not the vendor’s. Illinois BIPA, enacted in 2008, treats voiceprints as biometric identifiers, and a third-party vendor doing the collection does not shield the organisation that deployed it. The processing location is a compliance fact about you.

Then ask about the clock. Beamforming is a timing problem before it is a signal-processing problem, and channel skew of 50 microseconds corresponds to roughly 10.7 degrees of direction-of-arrival error at the geometry we calculated for. That is arithmetic setting a hardware requirement, again not a measurement: sound moves at 343 m/s, so the error follows from the spacing and nothing in software recovers it afterwards. A configuration of several USB microphones cannot satisfy it. The requirement is a single clock and a synchronous ADC on a common I2S or TDM bitclock.

Storage format is the third part of the same question. Lossy multichannel encoding does not preserve phase between channels, so an Opus or AAC multichannel file has already destroyed the spatial information that separation depends on. The answer you want is 48 kHz multichannel lossless, with a reference mono track alongside for anyone who wants to listen quickly. Ask, too, for the microphone array geometry as coordinates. Tools like ODAS require the physical positions of each capsule as configuration input, and a device that does not record its own geometry cannot have its beamforming reproduced later. If the vendor cannot separate two conversations at one counter, this section is where you will find out why.

Questions about the recording indicator

Ask whether the recording indicator can be turned off, by whom, and in which layer. This is the only question on the list with a direct legal consequence. California Penal Code section 632 makes it an offence to record a confidential communication without the consent of all parties, expressly including parties together in person.

Section 632 was amended by Stats. 2016 Ch. 855, effective 1 January 2017. A settings toggle in a phone app is not the same artefact as a firmware constant, and the difference is auditable. Our own answer is that colour and brightness are adjustable while the blink itself and a minimum brightness floor are hard-coded in firmware, which is a decision about what the product refuses to let anyone do. That is the argument behind a recording light with no off switch, and it is a narrower claim than a privacy policy because a policy can be rewritten by whoever writes policies.

Brewer v. Otter.ai, a putative class action filed in August 2025 in federal court in California, pleads ECPA, CFAA, CIPA and unfair-competition claims on the theory that the service acted as an unauthorised third-party eavesdropper. Whatever the outcome, the pleading is a description of the risk surface a notetaking device sits on, and the rules differ enough that recording consent varies by jurisdiction in ways a single global default cannot cover.

Questions about moving a day of data

Ask for the transfer interface, the measured throughput, and the resulting duration for one show day. A vendor who answers with an interface name and no rate has answered nothing, and the duration is the number that decides whether your staff go to dinner.

Do this division at the table. Bluetooth Low Energy at 2M PHY, best observed, carries about 0.7 Mbps. Thirty-two gigabytes over that link is about 101 hours, which is four days for one day of recording. The same 32 GB over Wi-Fi at a measured 50 Mbps is roughly 85 minutes, and over USB 2.0 mass storage at 35 MB/s about 15 minutes. These are quotients, not benchmark results, and any vendor can check them against their own numbers in a minute.

The consequence is architectural rather than a matter of preference. BLE is a control and status channel, which is how Open GoPro splits it, with Wi-Fi started by a BLE command. Bulk data goes over Wi-Fi or USB. A vendor still proposing Bluetooth as the primary data path for video has not run the arithmetic, and a proposal we ourselves wrote in an early requirements document was withdrawn for exactly this reason. Volume also sets the workflow: audio moves first because it is a fraction of the video, and the video goes overnight.

Questions about ownership and the moment power drops

Ask what survives an unplanned power loss, and ask who owns the calibration data. Exhibition halls lose power, someone trips over a cable, and a file being written at that instant either closes cleanly or does not. Ask for the segment length and what the recovery procedure is.

Segmented recording is what makes this answerable. Sixty-second segments mean a power loss costs at most one segment, and a segment is also a transfer checkpoint, so a resumed transfer requests missing segments rather than resuming a file at a byte offset. There is a trap in this area worth naming because we walked into it: the GStreamer splitmuxsink element has a max-files property whose documentation implies old files are deleted. We read the source. In the 1.24 branch, gstsplitmuxsink.c runs 4,158 lines and contains no call to g_remove or unlink at all; the property resets a fragment counter. With timestamped filenames nothing is ever reclaimed, no warning is raised, and the card fills to 100 percent. Any vendor claiming ring-buffer behaviour should be asked which code performs the deletion.

Calibration is the ownership question, and it is the quietest one on the list. Per-camera intrinsics and distortion coefficients determine whether footage can be rectified, stitched or measured later, and a unit shipped without them produces video that is difficult to reprocess outside the vendor’s own tooling. Ask for calibration values delivered per unit at shipment. Then ask the plain commercial version: what happens to your recordings if you stop paying. That is the same question as the note field that was always empty, asked about storage instead of about staff.

How to read the answer you get

A specific wrong answer is more useful than a confident vague one, because a specific answer can be checked and a vague one cannot. Listen for whether the vendor is quoting a datasheet maximum or a measurement they ran. The two sound similar and predict opposite outcomes.

Silicon marketing is where this shows most clearly. An RK3588 datasheet advertises 8K at 30fps, which is a single-stream maximum produced by committing two encoder cores to one stream, not headroom to divide across three cameras. An AM62A7 clears the compute requirement and then stalls on wiring, because it exposes one four-lane CSI-2 receiver, so three sensors need an external aggregator or SerDes, which is carrier-board work and bill-of-materials cost. Neither fact is a criticism of either part — both are reasons to require a measurement with three real sensors under sustained load rather than a synthetic benchmark.

Ask also what the device refuses to do. A product with three fixed profiles has decided something; a product with free combinations of resolution, framerate and bitrate has handed you a set of untested permutations and a way to lose a day by misconfiguring one. The same test applies to honesty about coverage, which is why we published what booth-fixed capture cannot do before anyone asked.

The last question is the one that sorts the field. Ask which of the previous answers the vendor will put in a document. We hold ourselves to the same standard, and it is the reason the capture sits in the counter rather than on a person: a claim that survives being written down is a different kind of claim. Everything else is a demo.

Primary sources for the legal points above: California Penal Code section 632, as amended by Stats. 2016 Ch. 855 (https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=PEN&sectionNum=632), and the Otter.ai class-action analysis published August 2025 (https://natlawreview.com/article/ai-notetaking-tools-under-fire-lessons-otterai-class-action-complaint).