#5319: Focus Stacking Without the Phone App Guesswork

Your phone app wants three shots. The optics want three hundred. Here's how to actually stack macro focus.

Featuring
Listen
0:00
0:00
Episode Details
Episode ID
MWP-5501
Published
Duration
34:37
Audio
Direct link
Pipeline
V5.2
TTS Engine
chatterbox-regular
Script Writing Agent
deepseek-v4-pro

AI-Generated Content: This podcast is created using AI personas. Please verify any important information independently.

Three focus points is a phone app convention, not an optical answer. At the magnifications used for macro work — five centimeter objects shot from a few millimeters away — depth of field collapses to a fraction of a millimeter. Spanning that object means a hundred or more focus slices. Three leaves enormous gaps, and the app produces something that looks fine on a phone screen while quietly blurring most of the detail you wanted.

The fix is to separate capture from processing. Shoot a bracket on whatever device you have, then hand the folder to a stacking engine on a desktop, a Linux box, or a container. The serious tools all work this way: Zerene Stacker, Helicon Focus, and the open source options. The algorithm scores each frame at each pixel position for local contrast — sharp areas have crisp edges, blurry areas don't — then picks the sharpest frame per location and stitches the winners together. Wavelet methods and depth-map methods are fancier variants of the same idea.

Frame count is where the real math lives. German macro photographer Daniel Knop's rule: every detail should be sharp in at least three frames, not one, so the software has overlap to make clean decisions. For a five centimeter object at two times magnification, that works out to three hundred to six hundred frames for full coverage. Nobody taps through that by hand — automated focus bracketing on a rail or tethered camera is the actual solution.

The open source landscape is well populated. focus-stack is a C++ library, MIT licensed, single OpenCV dependency, Dockerfile included, with OpenCL acceleration and roughly 100 MB of RAM per megapixel during merge. ChimpStackr offers a Python GUI and headless mode with four algorithms and raw support. shinestacker adds batch processing and a retouch editor. OpenFocus ships a pretrained deep learning fusion model. stackaroni is a single Rust binary with 16-bit TIFF support. There's no cloud SaaS for this — the cloud story is running these tools in containers, which fits neatly into an existing automation pipeline.

Rigging tolerances are more forgiving than you'd expect, up to a point. Stacking software registers frames before merging, correcting for scale changes, lateral drift, slight rotation, and exposure or white balance drift. ChimpStackr's author shot stacks of 150 frames at four times magnification on a "slightly wobbly rig" and got good results. But the documentation warns that deep stacks with very blurry extreme frames can fail to align, so the better the capture, the better the merge.

For shiny metal parts in a plastic container, take them out. Matte dark surface, clean edges, no reflections from the container walls. Then automate the bracket and let the software do the rest.

Downloads

Episode Audio

Download the full episode as an MP3 file

Download MP3
Transcript (TXT)

Plain text transcript file

Transcript (PDF)

Formatted PDF with styling

#5319: Focus Stacking Without the Phone App Guesswork

Corn
Daniel's back with a follow-up to the macro photography thread, and he's leaning into the pun this time. He wants to zoom in on focus stacking. His actual problem: he's photographing engraving bits, five or six centimeters long, from a few millimeters away, and the phone apps all seem to want exactly three shots, which he sets manually at three points along the object. He's asking whether three is some kind of magic number, what actually gives the best result for a row of metal carbides in a plastic container, and he's got a strong suspicion that the right answer is to shoot on whatever platform and do the stacking afterward in a separate tool, ideally something open source that runs on Linux or in a cloud environment. Then there's the rigging question. What kind of tolerances are we actually talking about if you want to do this properly.
Herman
Three shots is a phone app convention, not a physics answer. I want to be clear about that up front because it's the thing that's going to keep misleading him. At the magnifications he's working at, the depth of field is a fraction of a millimeter. A five centimeter object means you're spanning something like a hundred or more of those slices. Three focus planes leaves enormous gaps. The app isn't asking for three because three is right. It's asking for three because three is the number a person will tolerate tapping through before they get bored.
Corn
So the interface is setting the sample count, not the optics.
Herman
And Daniel's instinct to separate capture from processing is correct. The serious tools all work that way. Zerene Stacker, Helicon Focus, the open source ones. You shoot a bracket, you hand the folder to a stacking engine, it aligns and merges. The phone is a capture device. The desktop or the container is where the math happens.
Corn
Walk me through what the math is actually doing. If I hand a stacking program thirty images of the same engraving bit, what's it looking for?
Herman
For each pixel position, it's asking which frame has the sharpest version of that pixel. Sharpness, in this context, means local contrast. An out-of-focus area is blurry, so neighboring pixels are similar to each other. An in-focus area has crisp edges, high contrast between adjacent pixels. The algorithm scores each frame at each location and picks the winner. Then it stitches those winners into a single image. That's the basic depth-of-field extension method. There are fancier versions. Wavelet based methods decompose the image into frequency bands and merge the high frequency detail from whichever frame has it. And there are depth map methods that build a three dimensional model of the scene and project the sharp pixels onto it. But the core idea is the same. For every point on the object, some frame in the stack has it in focus. The software finds that frame.
Corn
And if no frame has it in focus, you get a soft spot.
Herman
You get a soft spot, or you get what the stacking community politely calls artifacts. Halos around edges, smeared textures, regions where the algorithm guessed wrong. That's where the frame count matters. Daniel Knop, the German macro photographer, has a rule. Every detail on the object should be sharp in at least three of the frames. Not one. Three. Because the software needs overlap between focus zones to make clean decisions. If each focus slice just barely touches the next one, the algorithm has no margin. It's guessing at the boundaries.
Corn
So three total shots is off by a factor of, what, fifty?
Herman
For a five centimeter object at a few millimeters working distance, I'd start at a hundred frames and see what the software does. Knop says five to ten frames is often enough for ordinary macro up to one to one magnification. But once you're beyond that, at five times or ten times magnification, he's using a hundred and fifty frames for a ten millimeter insect. Three hundred plus at twenty times. Daniel's engraving bits are not as demanding as a twenty to one insect eye, but they're a lot closer to that regime than they are to a portrait.
Corn
So the phone asking for three is like a car asking if you'd like to stop at the first, middle, or last gas station on a cross country drive. You're going to run out somewhere in Nebraska.
Herman
And the phone app will happily produce a stacked image from those three frames, and it'll look... fine. On a phone screen. Zoomed out. That's the trap. It looks like focus stacking worked, but what it actually did was pick the sharpest of three very thin slices and blur everything else into a smooth background. If Daniel wants to read the engraving on a carbide bit, or see the flute geometry, or check for burrs, he needs detail across the whole length. Three frames won't give him that.
Corn
Let's talk about the tool landscape then. He asked specifically for open source, Linux friendly, cloud possible. What's actually out there?
Herman
The strongest candidate is a C plus plus library called focus-stack. MIT licensed, single dependency on OpenCV, builds with make, and it ships a Dockerfile. That's the cloud angle. You throw it in a container, you mount a folder of images, you run it headless, you get a stacked output. It uses a complex wavelet algorithm, which is one of the better ones for this kind of work, and it has OpenCL acceleration if you've got a GPU available. Memory footprint is about a hundred megabytes per megapixel at default settings. So a twelve megapixel phone image is roughly one point two gigabytes of RAM during the merge. That's fine on a desktop, tight on a small VPS, trivial on a cloud instance with a few gigs.
Corn
A hundred megs per megapixel. That's the kind of number that tells you this is real computation, not a filter.
Herman
It's doing a wavelet transform across every frame simultaneously. It's not applying a sharpen slider. There's also a Python option called ChimpStackr, which has a graphical interface and a headless command line mode. It supports four different algorithms, including Laplacian pyramid and depth map. It handles raw files directly. The author's own gallery examples were shot at about four times magnification with roughly a hundred and fifty images per stack, on what he describes as a slightly wobbly rig. That's an important data point for the rigging question, by the way. We'll get there.
Corn
A hundred and fifty frames on a wobbly rig and the output still looks good?
Herman
That's the thing people don't appreciate about focus stacking software. Before it merges anything, it aligns the frames. Because when you refocus a lens, the image scale changes slightly. The lens elements move, the field of view shifts a tiny bit. And if you're moving the camera on a rail, there's going to be some lateral drift, maybe a fraction of a degree of rotation. The software corrects for all of that. Scale, translation, rotation, exposure drift, white balance drift. It's not just a merge, it's a registration pipeline.
Corn
So the software is forgiving of a certain amount of slop.
Herman
Up to a point. And that point matters for Daniel. The focus-stack documentation is explicit that deep stacks with very blurry extreme frames can fail to align. So the default behavior is to align each frame to its neighbor rather than to a single reference frame. That's a repair strategy for when the first and last frames are so different that they can't be matched directly. But it's still a repair. The better the capture, the better the merge.
Corn
What else is in the open source landscape? He asked for a survey, not just one recommendation.
Herman
There's shinestacker, Python, LGPL, Qt interface, batch processing for hundreds of images, and a retouch editor for fixing artifacts after the merge. That's a nice workflow touch. There's OpenFocus, which ships a pretrained deep learning fusion model called StackMFF version four. That's the new frontier. Neural network based fusion instead of classical wavelet or pyramid methods. There's a Rust option called stackaroni, single binary, sixteen bit TIFF support. Rust is nice because you compile it once and it runs anywhere without dependency hell. And there are a handful of smaller projects. StackAnt, Photo Focus Stacker, which came out of the OpenScan photogrammetry community. The field is not empty. It's actually quite well populated on the desktop and command line side.
Corn
And on Android, which is where he started his search?
Herman
That's the interesting part. The search wasn't wrong. Dedicated Android stacking apps are thin. But the engine already exists on Android. There's an app called Multifocus Camera that wraps the focus-stack library directly. Same algorithm, running on the phone. So the math is portable. What's missing is a good interface. The Android apps that do exist are mostly capture tools. BracketLab is open source, does raw focus and exposure bracketing through the Camera2 API. Macro Bee drives a motorized macro rail over Bluetooth and computes step size and shot count. Those are useful. But the actual stacking, the merge, the artifact cleanup, the retouching. That's a desktop job.
Corn
So Daniel's instinct is right. Shoot on the phone, stack on the Linux box. Or in a container.
Herman
And there's no cloud SaaS for this, by the way. I looked. There's no focus stacking web service where you upload a folder and get a merged image back. The cloud story is containers. You run focus-stack in a Docker container on whatever compute you have. That's actually a better fit for his workflow anyway. He's already running n8n. He could wire up a folder watcher that triggers a container build whenever he drops a new bracket in.
Corn
That's a compelling picture. Shoot a bracket on the phone, sync it to the server, the container chews on it, and the finished stack lands in his inventory system. No manual step in the middle.
Herman
And because focus-stack is a single binary with a single dependency, the container is small. You're not shipping a whole Python environment. You're shipping OpenCV and the tool. It's the kind of thing that runs fine on a five dollar VPS for occasional use, or on a cloud function if you're willing to deal with cold starts and memory limits.
Corn
Now the question he actually asked about the engraving bits. Five or six centimeters, row of metal carbides in a plastic container. Three manual focus points. What should he actually be doing?
Herman
The plastic container is the first problem. If the bits are in their container, the container walls are going to be in the frame. Depending on the angle, the plastic might be between the lens and the object, which adds reflections, or it might be behind the object, which adds clutter. Neither is ideal. For a serious stack, take them out of the container. Put them on a matte surface. Something dark, non-reflective. Metal carbides are shiny. They throw specular highlights everywhere. A dark background gives the stacking software clean edges to work with.
Corn
And the frame count?
Herman
Let's do the actual math. A five centimeter object at, say, two times magnification. Depth of field at that magnification, depending on aperture, is on the order of a few tenths of a millimeter. Let's call it half a millimeter to be generous. Fifty millimeters divided by half a millimeter is a hundred steps. But you need overlap. Knop's rule of three means you divide the step length by three to six. So you're looking at three hundred to six hundred frames if you want full coverage with proper overlap.
Corn
That's a lot of taps.
Herman
Which is why nobody does this by hand. The phone app's three shot workflow is a manual compromise. The real solution is automated focus bracketing. You set the near point, you set the far point, you set the step size, and the camera or the rail marches through the range. Helicon Remote does this over USB tethering. CamRanger does it wirelessly. On the phone side, Macro Bee computes the step size and drives a motorized rail. That's the hardware Daniel was asking about.
Corn
So the hardware exists, and it's not exotic. A motorized macro rail is a couple hundred dollars. The StackShot is the reference design.
Herman
StackShot is the professional standard. It's a stepper motor rail with sub-micron repeatability. You mount the camera on it, and it moves the whole camera between shots. That's actually the preferred method for high magnification work. You don't refocus the lens, because refocusing changes the image scale and complicates the alignment. You move the entire camera. The focus stays fixed, the lens stays fixed, and the rail advances by the step size. Knop does exactly this. Motorized linear stage, tethered monitor, moving the camera rather than the lens.
Corn
So the rigging question has a clear answer. Fixed camera, fixed lens, fixed focus, move the whole assembly on a rail. What kind of precision are we talking about?
Herman
For Daniel's use case, the step size is going to be somewhere in the range of a few tenths of a millimeter. Let's say half a millimeter steps with three times overlap means the rail needs to move about zero point one seven millimeters per shot, repeatably. That's not demanding by macro rail standards. A cheap stepper rail can do that. The StackShot does sub-micron. The cheap ones do maybe ten micron repeatability. That's fifty times better than he needs. So the precision question is almost a non-issue at his magnification.
Corn
What about the phone itself? He's shooting on a phone, not a DSLR on a rail.
Herman
That's where the rigging gets trickier. A phone has no tripod mount, no focus ring, and the lens is tiny. But the same principle applies. You want the phone fixed. A phone clamp on a tripod, or a small phone stand. And you want to move the phone, not refocus it. The problem is that phone cameras refocus automatically, and the focus is driven by software. You can lock focus in most camera apps by long pressing. But the focus step is not controllable in fine increments. You can't tell the phone to focus at exactly four point three millimeters. You can only tap on the screen and hope.
Corn
Which is why the manual three point method is so crude. He's tapping at three arbitrary points and the phone is doing whatever it wants in between.
Herman
Right. The phone's autofocus is not designed for this. It's designed to find a face or a subject and lock on. When you tap to focus on a tiny metal object a few millimeters away, the phone is hunting. It might lock at slightly different distances each time. That's why the alignment step in the stacking software is so important. It's compensating for the phone's focus sloppiness.
Corn
So the honest answer to his workflow question is: the phone is the weak link. Not the stacking software, not the rig, not the math. The capture device.
Herman
And there are two ways to fix that. One is to use a phone with a proper manual focus mode. Some Android phones expose manual focus through the Camera2 API. BracketLab uses it. You can set the focus distance in diopters and step through it programmatically. That's the right way to do focus bracketing on a phone. The other way is to use a dedicated camera with a macro lens and a tethered rail. That's the right way to do it if he wants to invest in hardware.
Corn
For an inventory system, though, the phone is the practical choice. He's photographing engraving bits, not museum specimens. The stakes are documentation, not publication.
Herman
And that changes the math. If the goal is to record what the bit looks like, so he can identify it later, he doesn't need three hundred frames. He needs enough frames to cover the critical detail. The engraving on the shank, the flute geometry, the tip. Maybe twenty or thirty frames, manually tapped, with the phone locked down. That's a perfectly reasonable inventory workflow. The three shot app default is still too few, but he doesn't need the full Knop treatment.
Corn
So there's a spectrum. Three shots is the phone app's guess. Twenty to thirty is the practical inventory bracket. Three hundred is the serious macro stack. And the tools scale across that whole range.
Herman
And the open source tools handle all of it. focus-stack doesn't care if you give it ten frames or five hundred. It just merges what you give it. The quality ceiling is set by the capture, not the software.
Corn
Let's talk about the alignment question one more time, because I think this is where Daniel's intuition about rigging is sharpest. He's presuming you want a perfectly consistent camera position. Is that actually true?
Herman
It's true in the sense that consistency makes everything easier. But the software is explicitly designed to handle inconsistency. The alignment step corrects for scale changes, translation, rotation, exposure drift. The focus-stack documentation says it aligns to compensate for camera movement. So a perfectly consistent position is not required. What's required is that the frames overlap enough that the alignment can find corresponding features. If the camera jumps around wildly, the alignment fails. If it drifts a few pixels between frames, the alignment handles it.
Corn
So the rigging tolerance is not zero. It's more like, keep it steady enough that the software can find its bearings. A cheap tripod and a phone clamp is probably enough.
Herman
For his use case, almost certainly. The ChimpStackr author's hundred and fifty frame stacks on a wobbly rig are the proof. If a slightly wobbly rig can produce good stacks at four times magnification, a phone on a desk tripod is going to be fine. The key is to minimize motion between frames. Not eliminate it entirely. The software is there to clean up the rest.
Corn
What about the exposure question? He's shooting shiny metal objects. The highlights are going to blow out. Does focus stacking help with that at all?
Herman
Not directly. Focus stacking is about depth of field, not dynamic range. But there's a related technique called exposure fusion, which some of these tools also support. ChimpStackr has a Mertens exposure fusion algorithm built in. That's for combining frames at different exposures, not different focus distances. If the metal bits are blowing out, he needs to handle that at capture time. Diffuse the light. Use a softbox or a piece of paper to bounce the light. Shiny metal is a lighting problem, not a focus problem.
Corn
And the plastic container he mentioned. That's also a lighting problem. Plastic reflects, and if the container has a clear lid, the lid is going to create a second surface that the stacking software has to deal with.
Herman
The lid is the worst case. If the bits are inside a clear plastic container with the lid on, the stacking software is going to try to focus on the lid surface, the bit surface, and the container bottom, all at different depths. It'll produce a mess. Take the bits out. Put them on a matte surface. Control the light. Then stack.
Corn
The workflow he should actually adopt is: remove the bits from the container, set up a small matte stage, clamp the phone, lock focus, tap through maybe twenty or thirty focus points along the length, and then hand the folder to focus-stack in a container. That's the whole pipeline.
Herman
If he wants to get fancy, he can script the focus stepping on the phone side. BracketLab can do programmatic focus bracketing through Camera2. That removes the manual tapping entirely. Set the near point, set the far point, set the number of steps, and the app captures the bracket. Then he syncs the folder and the container does the merge.
Corn
That's a practical answer. He asked for the landscape, and the landscape is: capture on the phone, stack on the Linux box, open source all the way down. The Android stacking apps are a dead end, but the Android capture apps are fine. The serious engines are all command line or desktop.
Herman
The cloud angle is solved by Docker. focus-stack ships a Dockerfile. That's the cloud story. Not a SaaS product, but a container you can run anywhere. For his setup, where he's already running self-hosted services, that's a better answer than a web service would be anyway.
Corn
I want to come back to the three shot question one more time, because I think there's a deeper point here. The phone app asks for three shots because three is the number of taps a human will tolerate. But the app doesn't explain that. It just presents three as the workflow. So Daniel, who's a careful person, assumed three was the standard. And it's not. It's a user interface compromise masquerading as a technical parameter.
Herman
That's the thing about phone camera software. It hides the real parameters behind a simplified interface. Three shots, auto mode, scene detection. The phone is making decisions for you and not telling you what they are. When you step into the world of focus stacking, you have to unlearn the phone's simplifications. Three is not a rule. It's a default. And the real rule is: enough frames to cover the depth with overlap.
Corn
The overlap rule, the three times minimum, that's the actual place where the number three appears. Not as a shot count, but as a coverage factor. Every detail sharp in at least three frames.
Herman
Which is a completely different concept. The phone app's three is a count. Knop's three is a redundancy factor. They happen to share a digit, but they're answering different questions. One is asking how many taps you'll tolerate. The other is asking how much overlap the software needs to make clean decisions.
Corn
If he's photographing a row of five or six carbides, six centimeters long, in a plastic container, the answer is: take them out of the container, light them diffusely, clamp the phone, and shoot a bracket with enough frames to cover the length with overlap. Twenty to thirty frames for inventory purposes. More if he wants publication quality.
Herman
Run the merge through focus-stack in a container. That's the reliable, effective, open source answer.
Corn
What about the focal length calculators he mentioned? He said he saw specific calculators for this application. Are those actually useful?
Herman
They are, but they're mostly aimed at DSLR and mirrorless users. Zerene Stacker has a published depth of field calculator that takes magnification, aperture, and sensor size, and gives you the step size. Macro by Raghu has a step size calculator for macro photography. The inputs are focal length, aperture, magnification, and circle of confusion. The output is the step size in millimeters. For a phone, the circle of confusion is tiny because the sensor is tiny, and the aperture is fixed, so the calculator is less useful. But the principle still applies. You want to know your depth of field so you know how many steps you need.
Corn
The circle of confusion numbers are interesting. The standard values are zero point zero two three millimeters for APS-C, zero point zero two nine for full frame, zero point zero one five for micro four thirds. So the smaller the sensor, the smaller the circle of confusion, which means the tighter the depth of field tolerance. A phone sensor is even smaller than micro four thirds.
Herman
Which is counterintuitive. People think small sensors have more depth of field, and they do, at the same field of view. But the circle of confusion is smaller, which means the acceptable sharpness threshold is tighter. The two effects partially cancel. It's one of those things where the simple story is wrong and the real story is more interesting.
Corn
The simple story being that phones don't need focus stacking because small sensors have deep depth of field. That's the iPhone macro argument. And it's true at normal distances. But at a few millimeters working distance, the depth of field collapses regardless of sensor size. The physics doesn't care how small your sensor is.
Herman
At a few millimeters, the depth of field is measured in tenths of a millimeter. A phone sensor doesn't save you from that. It just means the circle of confusion is smaller, so the acceptable blur is tighter. You still need stacking. The phone's deep depth of field advantage only applies at normal subject distances.
Corn
The people who say phones don't need focus stacking are technically right, but only in the regime where depth of field is already large. Once you're doing macro, the phone is just as limited as a DSLR. More limited, actually, because you have less control.
Herman
That's the real takeaway for Daniel. The phone is a perfectly good macro camera in some ways. The small sensor gives you deep depth of field at moderate distances. But for the kind of work he's doing, a few millimeters from a metal object, the phone's limitations show up. No fine focus control, no aperture control, no tripod mount. The stacking software can compensate for some of that, but the capture is the bottleneck.
Corn
The answer to his question, what's the best, most effective, reliable way to do this, is a pipeline. Phone capture with locked focus and a stable mount. Twenty to thirty frames for inventory work. Open source stacking engine in a container. Matte surface, controlled light. That's the whole thing.
Herman
The open source landscape is good. Focus-stack for the command line and container work. ChimpStackr if he wants a GUI and raw support. Shinestacker if he wants batch processing and retouching. OpenFocus if he wants to play with the deep learning fusion models. The field is not empty. It's just not on the Android app store.
Corn
Which is the answer to the question he didn't quite ask, but was circling. Why did the Android app store come up short? Because the serious tools are not phone apps. They're desktop and command line programs. The phone is a capture device. The stacking engine lives somewhere else.
Herman
The one Android app that does real stacking, Multifocus Camera, is just a wrapper around the same focus-stack library he'd be running on his Linux box. So the math is portable. The interface is not. That's the gap.
Corn
Let's talk about the deep learning angle for a minute, because that's the part of the landscape that's actually moving. OpenFocus ships a pretrained neural fusion model. There's a paper from last December about something called Fovea Stacking, which uses deformable phase plates in the optics to do the stacking physically rather than computationally. That's a hardware approach that outperforms traditional focus stacking in image quality.
Herman
The Fovea Stacking paper is interesting. It's an ACM Transactions on Graphics paper, which is the top venue for this kind of work. The idea is that instead of taking a hundred frames and merging them in software, you put a deformable phase plate in the optical path that changes the focus during a single exposure. The sensor captures a single image that already contains the depth information. Then you reconstruct the sharp image from that single capture.
Corn
It's focus stacking in a single shot, using hardware to encode the depth information.
Herman
That's the pitch. And the results are better than traditional stacking in image quality, according to the paper. But it's research hardware. Deformable phase plates are not something you buy on AliExpress. This is a lab technique. The practical takeaway for Daniel is that the field is evolving, and the direction is toward more computational and optical sophistication, not less.
Corn
Which means the open source tools he can use today are going to get better, not worse. The wavelet algorithms in focus-stack are already excellent. The deep learning models in OpenFocus are the next generation. And the hardware approaches like Fovea Stacking are the generation after that.
Herman
The nice thing about the open source landscape is that it tracks all of this. focus-stack is maintained. ChimpStackr is maintained. OpenFocus is actively developing new fusion models. He's not betting on a dead ecosystem. He's betting on a live one.
Corn
Let's circle back to the practical question one more time, because I want to make sure we actually answered what he asked. He was photographing engraving bits from AliExpress. Five or six centimeters long, in a plastic container. He manually set three focal points. What should he do differently tomorrow?
Herman
Tomorrow, he should take the bits out of the container. Put them on a dark matte surface. Clamp the phone to something stable. Lock focus manually. And then instead of three taps, do twenty or thirty taps along the length of the object. Twenty or thirty. Then run the folder through focus-stack on his Linux machine or in a container. That will give him a clean, sharp image of the entire bit, from tip to shank, with the engraving legible and the flute geometry visible.
Corn
If he wants to automate the capture, he should look at BracketLab for programmatic focus bracketing on the phone.
Herman
BracketLab is open source, uses the Camera2 API, and can do raw focus and exposure bracketing. It's not a stacking app. It's a capture app. But that's exactly what he needs on the phone side. Capture the bracket, then hand it to the stacking engine.
Corn
The workflow is: capture on the phone with a real bracketing app, stack on the Linux box with an open source engine, and the whole thing can run in a container if he wants to automate it. That's the answer.
Herman
The rigging is simpler than he thinks. A phone clamp on a cheap tripod is enough. The alignment software handles the rest. He doesn't need a StackShot rail for inventory photos of engraving bits. He needs a stable phone and enough frames.
Corn
The StackShot is for people who are stacking three hundred frames at twenty times magnification. Daniel is stacking twenty frames at two times magnification. The precision requirements are orders of magnitude apart.
Herman
That's the thing about precision tolerances. They scale with magnification. At two times, a ten micron wobble is nothing. At twenty times, it's a disaster. Daniel is nowhere near the regime where the expensive rails are necessary.
Corn
The answer to his rigging question is: don't overthink it. A stable phone mount is sufficient. The software will absorb the rest.
Herman
The answer to his three shot question is: three is a phone app convention, not a physics answer. The real rule is overlap. Every detail sharp in at least three frames. For a five centimeter object, that means twenty to thirty frames for inventory work, more for publication quality.
Corn
Which is a much more satisfying answer than the app store gave him.

Hilbert: The focus-stack Dockerfile pulls Ubuntu with the OpenCV development package. I know because I built it once for a batch of circuit board photos. The container came out at just under two gigabytes. Ran fine on a four gig instance. The merge took about three minutes for forty frames at twelve megapixels.
Corn
That's a useful data point. Three minutes for forty frames is entirely reasonable for an inventory workflow.

Hilbert: The memory is the thing to watch. A hundred megabytes per megapixel means a twelve megapixel image wants one point two gigs. If you run two merges at once on a small box, you're swapping. Run them one at a time.
Herman
That matches the documentation. There's a flag to reduce the batch size and thread count, which brings it down to about fifty megabytes per megapixel. Slower, but it fits in a smaller container.

Hilbert: I ran it with a single thread and a batch size of two. Took twice as long, but it stayed under seven hundred megs. Good enough for a five dollar VPS.
Corn
The cloud story is real. You can run this on the cheapest tier if you're patient.

Hilbert: The alignment is the part that surprised me. I had a few frames where the board shifted because I bumped the desk. The output still came out clean. It found the edges and corrected for the shift. I wouldn't have believed it if I hadn't seen the before and after.
Herman
That's the neighbor-to-neighbor alignment. It's designed for exactly that. Small shifts between frames get corrected. Large jumps break it, but small bumps are fine.

Hilbert: The other thing. Don't put the bits back in the plastic container before you shoot. The container lid adds a reflection layer that the merge can't fix. I learned that the hard way with some small screws.
Corn
The workflow advice holds. Remove from container, matte surface, stable phone, twenty to thirty frames.

Hilbert: If you're going to do it regularly, write a script. The whole thing is three commands. Sync the folder, run the container, copy the output. You could do it with a cron job.
Herman
That's the n8n integration Daniel was probably already thinking about. Folder watcher triggers the container, container produces the stack, stack goes into the inventory system. No manual step.
Corn
The one thing I'd want him to remember from all of this is that three is not a rule. It's a default. The phone app asks for three shots because three is what a person will tolerate, not because three is what the optics require. The actual requirement is overlap. Every detail sharp in at least three frames. That's the real three.
Herman
The corollary. The phone is the weak link, not the software. The open source stacking engines are excellent. The bottleneck is capture. Fix the capture, and the rest of the pipeline is solved.
Corn
Which is a nice place to land. Daniel's instinct to separate capture from processing was right. The Linux and cloud tools are there. The rigging is simpler than he feared. And the answer to his AliExpress engraving bits is: take them out of the container, light them diffusely, clamp the phone, and shoot twenty to thirty frames instead of three.
Herman
That's the episode. Thanks to Hilbert Flumingtop for producing, and for the container memory numbers.
Corn
This has been My Weird Prompts. If you want to send us a question, email us at show at my weird prompts dot com.
Herman
We'll be back soon.

This episode was generated with AI assistance. Hosts Herman and Corn are AI personalities.