← Firmware research← 韌體研究

In-camera lossless compression機身內的無損壓縮

The fp has a hardware lossless-JPEG engine. It compresses every still you take and it has never once touched a video frame. The engine works, the format is legal, the ratio is enough — what is tight is milliseconds per frame, and that number has now been measured.

fp 裡有一顆硬體的無損 JPEG 引擎。你拍的每一張照片都經過它, 而它從來沒有碰過任何一格影片。引擎是好的、格式是合法的、壓縮比也夠 —— 緊的是每一格有幾毫秒,而這個數字現在量出來了。

What is measured量到的是什麼
The engine's full cost function, taken on the camera by wrapping the stills path's own encode call. Three terms, fitted to four points, residual within ±2 µs. Compression ratios come from the engine's own size table, not from file sizes.
引擎完整的成本函數,是在相機上把拍照路徑自己的編碼呼叫包起來量的。 三個項、用四個點擬合,殘差在 ±2 µs 以內。 壓縮比來自引擎自己的尺寸表,不是從檔案大小反推的。
What is not沒量到的是什麼
Nothing end-to-end. No compressed video frame has been written to a card by the camera. The numbers below are a budget, not a recording.
沒有任何一條完整走通的路。相機從來沒有把一格壓縮過的影片寫進卡裡。 下面的數字是預算,不是錄影結果。
The blocking question卡住的問題
Whether the camera can play back what it would write. The movie parser does not read the tile tags, so a compressed clip may be export-only — see dynamic frame-skip.
相機能不能播放它自己寫出來的東西。影片解析器不讀 tile 標籤, 所以壓縮過的片段可能只能匯出、不能在機上播 —— 見動態抽幀。

Where compression sits壓縮在哪一關

Before any of the arithmetic, the one thing that decides what you are allowed to plan: compression happens last.

在算任何數字之前,先記住決定你「能規劃什麼」的那件事: 壓縮發生在最後。

The pipeline, with compression after all geometry is fixed sensor readout 感光元件讀出 reduction — ONE branch of an 8-way selector, not both in series 縮減 — 八選一的其中一條,不是兩個串起來 Hbin2  36 profiles Hbin2  36 個 profile RWZM  53 profiles RWZM  53 個 profile DRAM strip DRAM strip packed 12-bit CinemaDNG 打包的 12-bit CinemaDNG all reduction happens here — 253 of 306 profiles reduce nothing at all 所有縮減都在這裡 — 306 個 profile 裡有 253 個完全不縮 0x300D lossless JPEG 0x300D 無損 JPEG compression happens here 壓縮發生在這裡 SD / SSD no way back 回不去了 after this the data is a bit-stream, 在這之後資料是位元流, not pixels 不是像素
The engine reads the finished CinemaDNG strip — the same buffer the writer is about to put on the card. Once a frame is compressed you cannot hbin, vbin or RWZM it.
引擎讀的是已經做完的 CinemaDNG strip —— 就是 writer 正要寫進卡裡的那塊緩衝區。 一格畫面一旦被壓縮,就不能再 hbin、vbin 或 RWZM 了。
Decide geometry first, then talk about compression. Three reasons, any one of which is enough. Hardware order: binning and resampling are branches of the CRCT pipeline — a profile takes one of them, not both in series — and all of them sit physically before the write to DRAM; the engine is after it. (The branch census: 253 profiles reduce nothing, 36 take Hbin2, 17 take RWZM together with hbin.) Data type: the engine's output is a Huffman bit-stream, not pixels — any reduction would need a full decode first, and the camera's video path has no such decode. Sizing: compression changes bytes, never geometry. Work out the frame size in the reduction path first; a ratio computed on a frame size you have not settled is a ratio for a frame that does not exist.
先把幾何決定好,再談壓縮。三個理由,任何一個都足夠。
硬體順序:binning 與重取樣是 CRCT 管線的分支 —— 一個 profile 只會走其中一條,不是兩個串起來 —— 而且它們全部在寫進 DRAM 之前,引擎在那之後。 (分支普查:253 個 profile 什麼都不縮、 36 個走 Hbin2、17 個 RWZM 與 hbin 同時。)
資料型態:引擎的輸出是 Huffman 位元流,不是像素 —— 要再縮就得先完整解碼,而相機的影片路徑沒有這種解碼器。
尺寸:壓縮改變的是位元組,從來不改幾何。 先去縮減路徑那頁把畫面尺寸定下來; 在一個還沒定案的尺寸上算出來的壓縮比,是一個不存在的畫面的壓縮比。

What the engine costs引擎的成本

Measured on the camera, by wrapping the stills path's own encode call rather than cold-calling the engine. Three terms, and they genuinely separate:

在相機上量的,做法是把拍照路徑自己的編碼呼叫包起來, 而不是冷呼叫引擎。三個項,而且是真的可以拆開的:

The three-term cost model and how it splits for each format t(µs) = 2,179 + 4,954.40 × alignedMpix + 40.926 × tiles per call 每次呼叫 201.8 Mpix/s per tile 每個 tile FHD OG3K UHD 6K 13.87 ms 34.33 ms 44.91 ms 130.8 ms call 呼叫 pixels 像素 tiles tile
The per-call term is a flat 2.18 ms whatever you encode, so it hurts small frames most — it is 16% of an FHD frame and 1.7% of a 6K one. The per-tile term is why fewer, larger tiles beat less padding.
「每次呼叫」那一項不管你編什麼都是固定 2.18 ms, 所以對小畫面傷害最大 —— 它佔 FHD 一格的 16%,佔 6K 一格只有 1.7%。 而「每個 tile」那一項,就是為什麼用比較少、比較大的 tile,勝過減少 padding。
Two older numbers you will find in the notes are wrong or misleading. 32 Mpix/s was a watchdog timeout, never a throughput. 169.7 Mpix/s is not wrong but is an effective rate at 256×256 that swallows the per-tile term — splitting it out gives 201.8 Mpix/s plus the other two terms. Always budget with all three.
筆記裡有兩個舊數字,是錯的或會誤導。 32 Mpix/s 是一個看門狗逾時,從來不是吞吐量。 169.7 Mpix/s 不算錯,但它是 256×256 之下的有效速率, 把「每個 tile」那一項吃進去了 —— 拆開來是 201.8 Mpix/s 再加另外兩項。 算預算時三項都要算。

Tile geometry is worth up to a quarter of the timeTile 幾何最多值四分之一的時間

The simulator searches this rather than looking it up, over the engine's own limits — tile width 32…512 in steps of 32, tile height even and at most 512, at most 326 tiles — plus TIFF's requirement that TileLength be a multiple of 16. That last constraint matters: without it the search picks 512×364 for FHD, which is genuinely faster and not a valid tag. With it, the search reproduces all four measured optima below exactly.

模擬器是搜尋出來的,不是查表,搜尋範圍就是引擎自己的限制 —— tile 寬 32…512(32 的倍數)、tile 高是偶數且最多 512、最多 326 個 tile —— 再加上 TIFF 要求 TileLength 必須是 16 的倍數。
最後這個限制很關鍵:沒有它的話,搜尋會替 FHD 挑出 512×364 —— 那確實比較快,但不是合法的標籤值。加上它之後, 搜尋結果與下面四個實測最佳值完全一致。

format格式best tile最佳 tiletilestile 數encode編碼vs 256×256對比 256×256
FHD 1936×1090512×3681213.87 ms−17.4%
OG3K 3024×2010512×5122434.33 ms−7.9%
UHD 3840×2160480×4324044.91 ms−12.9%
6K 6064×4042512×51296130.8 ms−8.3%

Bit depth: 8-bit is a different strategy, not an extra one位深:8-bit 是另一條路,不是多一個選項

The codec and the raw path do not agree about what a sample is. The engine's setup function maps a bitdepth code to a width and refuses anything it does not recognise:

codec 跟 raw 路徑對「一個取樣有多寬」的看法不一致。 引擎的設定函式把位深代碼對應到寬度,認不得的一律拒絕:

code代碼0123anything else其他任何值
engine sample width引擎的取樣寬度12141610 rejected — the call returns failure直接拒絕 — 呼叫回傳失敗

The raw output path does have 8-bit — its own depth field carries 8, 10, 12 and 14, and the debug print has a branch for each. So 8-bit frames are something the camera can produce, and something the compressor will not touch.

raw 輸出那條路確實有 8-bit —— 它自己的位深欄位帶 8、10、12、14,而且除錯輸出每一種都有分支。 所以 8-bit 的畫面是相機做得出來、但壓縮器不會碰的東西。

Ways to get FHD 29.97 under a card's sustained rate 12-bit, uncompressed 12-bit,未壓縮 10-bit, uncompressed 10-bit,未壓縮 8-bit, uncompressed 8-bit,未壓縮 12-bit at 2.0:1 12-bit,2.0:1 12-bit at 2.525:1 12-bit,2.525:1 97.2 MB/s 81.4 65.6 49.8 39.9 SD 94 MB/s FHD 1936×1090 at 29.97, including the 78,848-byte header each frame carries FHD 1936×1090 @29.97,含每格 78,848 bytes 的標頭
Dropping to 8-bit costs a third of the data and no engine time at all. Compressing at 2:1 saves more and is lossless, but costs 13.87 ms a frame and only works at 10, 12, 14 or 16 bits. You cannot have both.
降到 8-bit 可以省掉三分之一的資料量,而且完全不花引擎時間。 壓到 2:1 省更多而且是無損的,但每一格要花 13.87 ms, 而且只有 10、12、14、16 bit 能用。兩者不可兼得。

That makes 8-bit the cheap answer to a bandwidth problem and compression the expensive one. 8-bit is lossy and instant; lossless compression keeps every bit and costs a large slice of the frame budget. Picking 8-bit in the simulator therefore switches compression off rather than stacking the two — which is what the hardware would do.

所以 8-bit 是頻寬問題的便宜解,壓縮是昂貴解。 8-bit 有損但立即;無損壓縮保留每一個位元,但吃掉一格預算裡很大一塊。 因此在模擬器裡選 8-bit 會把壓縮關掉,而不是兩者疊加 —— 硬體本來就是這樣。

How much it actually compresses實際上壓得掉多少

Read out of the engine's own size table. It moves with the picture, and the spread is wide enough that it changes conclusions:

這些是從引擎自己的尺寸表讀出來的。它會隨畫面內容變動, 而且變動幅度大到足以改變結論:

Compression ratio range 1.5 1.65 2.0 2.6 2.8 stress floor 最差情況下限 12-bit design value 12-bit 設計值 dark-scene measured 暗場景實測 14-bit stills, effective samples: 1.67–2.08 · 12-bit video frames measured so far: 2.51–2.74 (all one dark scene) 14-bit 照片的有效取樣:1.67–2.08 · 目前量到的 12-bit 影片格:2.51–2.74(全部來自同一個暗場景)
Plan against 1.65:1, design against 2.0:1, and treat 2.6:1 as what a friendly scene gives you. The high-ISO-at-correct-exposure curve has still not been measured, because shutter and aperture cannot be written from the host.
規劃時用 1.65:1,設計時用 2.0:1, 2.6:1 當成「畫面很配合時的運氣」。
「高 ISO 但曝光正確」那條曲線還沒量到,因為快門和光圈沒辦法從主機端寫入。
Do not derive a ratio from DNG file sizes. A full still is mostly not raw tiles — one 36.33 MB file holds 25.68 MB of tiles and about 10.65 MB of other IFDs and previews. Ratios here come from summing TileByteCounts, which is the engine's own accounting.
不要從 DNG 的檔案大小去反推壓縮比。 一張完整的照片裡大部分不是 raw tile —— 一個 36.33 MB 的檔案裡,tile 佔 25.68 MB, 另外約 10.65 MB 是其他 IFD 和預覽圖。
這裡的壓縮比是把 TileByteCounts 加總出來的,那是引擎自己的帳。

The architecture is serial — encode and write add up架構是串列的 — 編碼和寫入要相加

This is the constraint that decides more than the engine speed does. In the current hook, one worker blocks on the encode, then blocks on the SD flush, and only then takes the next frame:

這個限制比引擎速度更決定成敗。在目前的 hook 裡, 一個 worker 先卡在編碼上,再卡在 SD 的 flush 上,然後才去拿下一格:

Serial encode-then-write versus a pipelined arrangement today — one worker, serial 目前 — 一個 worker,串列 encode 13.9 編碼 13.9 write 13.4 寫入 13.4 frame budget 33.4 ms @30p 一格預算 33.4 ms @30p service time 27.3 ms — fits, but both terms count 服務時間 27.3 ms — 夠,但兩項都要算 OG3K at the same scale OG3K,同一個比例尺 encode 34.3 編碼 34.3 service time 70.0 ms — over budget before the write even finishes 服務時間 70.0 ms — 還沒寫完就已經超出預算
Engine duty and card speed are not two resources that overlap in the current design. Separating them into two workers with a queue is a real project, not a tuning knob — and it is what OG3K would need.
在目前的設計下,引擎佔用與卡片速度不是兩個會重疊的資源。 把它們拆成兩個 worker 加一個隊列是一個真正的專案,不是一個調校旋鈕 —— 而 OG3K 需要的正是這個。

Budget simulator預算模擬器

Everything above, wired together. The cost function only reads the frame size, so the list is the camera's 13 sensor rasters (from the mode table at 0xC0B59DC0) rather than all 70 modes — the mode id never changes a number here. Pick a raster and a frame rate independently, say how much is reduced before the codec sees it, and the tile geometry is searched rather than guessed. “Custom size” at the bottom of the list takes any width and height, so you can cost a frame the camera does not currently have — it is checked against the engine's real limits (width 8…16384 in steps of 8, height 2…16384 even, at most 326 tiles) and told to you when it falls outside them.

把上面所有東西接起來。成本函數只讀畫面尺寸, 所以清單是相機的 13 個感光元件光柵(來自 0xC0B59DC0 的模式表), 而不是全部 70 個模式 —— mode id 在這裡不會改變任何一個數字。
光柵和格率各自獨立選,再指定進 codec 之前縮了多少, tile 幾何是搜尋出來的而不是猜的。
清單最後的「custom size」可以輸入任意寬高, 所以你能替相機目前沒有的畫面算成本 —— 輸入會對照引擎的真實限制檢查(寬 8…16384 且是 8 的倍數、 高 2…16384 且是偶數、最多 326 個 tile),超出範圍會直接告訴你。

Two things about those numbers. Frame rate is yours to choose, not tied to the raster: every rate the mode table carries is in the list, including the exact NTSC variants (a mode that advertises 30 runs at 29.97, one that advertises 78 runs at 77.27), and the detail line tells you which rates stock firmware actually offers at the raster you picked. And the reduction is plain arithmetic on the raster — the real path also crops, which is why 3032×1708 at 1.5625× lands on 1936×1092 here against the 1936×1090 the camera writes.

關於這些數字有兩件事。格率由你選,不綁在光柵上: 模式表帶的每一個格率都在清單裡,包含精確的 NTSC 變體 (標稱 30 的模式實跑 29.97、標稱 78 的實跑 77.27), 而底下的明細行會告訴你原廠在你選的那個光柵上實際提供哪些格率。
另外,縮減在這裡是對光柵做單純的算術 —— 真實路徑還會裁切, 所以 3032×1708 在 1.5625× 之下這裡算出 1936×1092, 而相機實際寫的是 1936×1090。

Start with the compression switch. Turn it off and every number becomes the raw rate, so what compression buys is the difference between the two. The other three toggles only shape how it compresses, and at a setting where the engine has time to spare none of them move the total — that is correct, not a fault.

先動 compression 那個開關。 把它關掉,每個數字都會變成未壓縮的碼率,兩者的差就是壓縮實際換到的東西。
另外三個開關只影響怎麼壓;在一個引擎還有餘裕的設定下, 它們都不會改變 total —— 那是正確的,不是壞掉。

Frame畫面

Compression壓縮

2.00 : 1

Result結果

—
engine can process引擎處理得了
—MB/s
from compressed frames來自壓縮的幀
—MB/s
from uncompressed frames來自沒壓的幀
—MB/s
if nothing compressed完全不壓的話
—MB/s
total to card寫到卡的總量
—MB/s
compressed payload壓縮後的資料 uncompressed payload沒壓的資料 headers, 78,848 B/frame標頭,78,848 B/格
encode編碼 — write寫入 — serial total串列總和 — budget一格預算 —

Dynamic frame-skip: legal, but not free動態抽幀:合法,但不是免費的

Compress the frames the engine can keep up with, write the rest uncompressed. The format allows it outright — CinemaDNG is one file per frame and Compression is a per-IFD tag, so a clip with a mix of both is valid and no frame is dropped. Three things it does not solve:

引擎跟得上的那些幀就壓,其餘的直接寫未壓縮。 格式本身完全允許 —— CinemaDNG 是一格一個檔, 而 Compression 是每個 IFD 各自的標籤, 所以一段片子裡兩種混著存是合法的,而且不會掉任何一格。 但它有三件事解決不了:

  1. It lowers the average, not the peak它降的是平均,不是尖峰

    The uncompressed frames still arrive at full rate. A card fails on sustained throughput, so “25% less on average” is not the same as “the clip that used to drop frames now does not.” Turn on partial-frame mode in the simulator to see the peak move instead.

    沒被壓的那些幀還是整張全速送過來。 卡是在持續吞吐上垮掉的,所以「平均少了 25%」不等於 「本來會掉幀的片子現在不掉了」。
    在模擬器裡打開「部分畫面模式」,就會看到尖峰跟著動。

  2. The camera probably cannot play it back相機大概播不出來

    The movie parser reads StripOffsets and never looks at TileOffsets, TileWidth or Compression. The stills path has a hardware LJPEG branch; the video path does not go through it. A compressed clip may be export-only.

    影片解析器只讀 StripOffsets, 從來不看 TileOffsets、TileWidth 或 Compression。拍照路徑有硬體 LJPEG 的分支;影片路徑不走那裡。 所以壓縮過的片子可能只能匯出,不能在機上播。

  3. The timeline breaks in three different ways時間軸有三種壞法

    Number the files consecutively and the skipped time vanishes — motion speeds up. Keep the original numbers and you have gaps, and nobody has checked what the player does with those. Repeat the previous frame and you keep the timing but lose the I/O saving that was the point. Real variable-rate needs per-frame timestamps or a sidecar, and something that understands them.

    檔名連號編下去,被跳過的時間就消失了 —— 動作會變快。
    保留原本的編號就會有空號,而沒有人檢查過播放器拿到空號會怎樣。
    重複前一格可以保住時間,但就失去了當初要省的 I/O。
    真正的可變格率需要逐格時間戳或 sidecar,以及看得懂它們的東西。

So frame-skip cannot be hidden as an implementation detail. The three safe product shapes are: drop the whole clip one frame-rate step, mark compressed clips export-only and ship a sidecar timeline, or build a playback layer that understands variable rate. Until one of those exists this stays a research probe.

所以抽幀不能當成一個「實作細節」藏起來。 三個安全的產品形態是:整段片子降一階格率、 把壓縮過的片子標成只能匯出並附一個 sidecar 時間軸、 或者做一層看得懂可變格率的播放。
在這三者之一出現之前,這件事只是一個研究性的試探。

Partial-frame mode部分畫面模式

The simulator's third toggle compresses a fraction of each frame instead of a fraction of the frames. Arithmetically it turns the compressed fraction from a whole-frame ratchet into a continuous dial, and the useful consequence is that every frame ends up the same size — so the peak drops with the average instead of staying pinned at the uncompressed rate.

模擬器的第三個開關,壓的是每一格的一部分,而不是一部分的格。 算術上,它把「壓縮比例」從整格跳動的棘輪變成連續的旋鈕; 有用的後果是 每一格最後大小都一樣 —— 所以尖峰會跟著平均一起降,而不是一直卡在未壓縮的碼率上。

It is not a legal DNG. An IFD is either tiled Compression=7 or an uncompressed strip; there is no half-and-half encoding. Doing it for real means a custom container or a sidecar, which lands you right back in the playback problem above. The toggle is there to make the average-versus-peak difference visible, not because it can be shipped.
它不是合法的 DNG。一個 IFD 要嘛是 tile 化的 Compression=7,要嘛是未壓縮的 strip,沒有一半一半的編碼方式。 真要做就得自訂容器或 sidecar,那又掉回上面那個播放問題。
這個開關存在的目的,是讓「平均 vs 尖峰」的差別看得見,不是因為它能出貨。

Verdict as it stands目前的結論

target目標why理由
FHD 23.976 → SDyes可以 27.3 ms serial against a 41.7 ms budget. The first thing to try.串列 27.3 ms,預算 41.7 ms。第一個該試的。
FHD 29.97 → SDlikely應該可以 Same 27.3 ms against 33.4 ms. Fine on a 94 MB/s card, marginal at 66.一樣 27.3 ms,對 33.4 ms。94 MB/s 的卡沒問題,66 就很勉強。
OG3K 30p → SDno Engine 3% over (20% with contention), card 18% over. Both need solving and the engine has no clock knob.引擎超 3%(算爭用是 20%),卡超 18%。兩個都要解,而且引擎沒有時脈旋鈕。
OG3K 30p → SSDnot needed不需要 Uncompressed 275.6 MB/s already fits in 390. Compression buys recording time, not feasibility.未壓縮的 275.6 MB/s 本來就塞得進 390。壓縮買到的是錄影時間,不是可行性。
OG3K 18/24 → SDno 70–85 ms per frame serial. Needs a real pipeline, not a tile change.串列每格 70–85 ms。需要的是真正的管線,不是換 tile。
UHD, any rateUHD,任何格率no Engine-bound at 44.91 ms. Nothing to do with the ratio.卡在引擎的 44.91 ms。跟壓縮比無關。
6Kno 314% duty at 24p.24p 之下負載 314%。

Still open還沒解的

Sources出處

The cost model, tile study and ratio measurements are from the in-camera lossless work of 2026-09-15/16; the hardware inventory, engine ABI and cold-call analysis are from the raw-compression research of 2026-08-20; the pipeline position comes from the reduction-path work. A summary of all of it, with the numbers this page uses, is in research/imaging-hw/notes/LOSSLESS_DEVELOPMENT_SUMMARY.md in the repository.

成本模型、tile 研究與壓縮比量測來自 2026-09-15/16 的機身內無損壓縮工作; 硬體清單、引擎 ABI 與冷呼叫分析來自 2026-08-20 的 raw 壓縮研究; 管線位置來自縮減路徑那邊的工作。
全部的彙整、以及這一頁用到的數字,在程式庫的 research/imaging-hw/notes/LOSSLESS_DEVELOPMENT_SUMMARY.md。