RawCull Tech Documentation
RawCull Developer Notes
These pages explain RawCull from a developer’s point of view. They are intended
for readers who already understand Swift and SwiftUI and now need to learn where
behavior lives, how data moves through the app, and which invariants must
survive a change.
Source paths in these pages are relative to the named repository. App paths such
as RawCull/Model/ViewModels/RawCullViewModel.swift belong to the RawCull
repository. Package paths belong to the separately versioned Swift packages
referenced by RawCull.xcodeproj; this documentation repository does not
contain or compile source snapshots.
Start Here
Read these pages in this order when learning the project:
- Thumbnails and Scan Pipeline follows a selected catalog from
directory discovery to visible images.
- Concurrency explains
@MainActor, actor ownership,
cancellation, and bounded parallelism across that pipeline. - Cache System explains why grid, preview, disk, full-size, and
analysis caches are separate.
- Focus Mask and Sharpness follows scoring and focus evidence
into persisted culling data.
- Burst Groups combines sharpness, similarity artifacts,
grouping, ranking, and the burst-review UI.
- Artificial Intelligence and RawCull Packages explain the
reusable package boundaries behind the app.
Use File Read and Write and
Security-Scoped URLs whenever a change touches persistence, caches,
user-selected folders, export, or copying.
Repository And Target Map
| Area | Primary source | Responsibility |
|---|
| App composition | RawCull/Main/RawCullApp.swift | Creates the AI composition root and shared observable models, injects environment state, and defines app windows and commands |
| Main presentation | RawCull/Main/RawCullMainView.swift | Switches between loupe, grid, similarity, rated, and comparison modes and owns top-level sheets and overlays |
| App orchestration | RawCull/Model/ViewModels/RawCullViewModel.swift plus feature extensions | Owns main-actor catalog, selection, navigation, progress, culling, similarity, and burst-review state |
| Background ownership | RawCull/Actors/ | Serializes scans, thumbnail requests, caches, persistence, extraction, and contention gates |
| App services and adapters | RawCull/Model/ | Connects the UI model to RAW parsing, sharpness analysis, AI providers, persistence, diagnostics, and rsync |
| Reusable domain logic | RawCullCore package | Value models and pure algorithms such as burst grouping and ranking |
| RAW decoding | RawParserKit package | ARW, NEF, and DNG dispatch, metadata normalization, MakerNote parsing, thumbnail extraction, and preview creation |
| Image analysis | PhotoAnalysisKit package | Sharpness, saliency, focus evidence, masks, and analysis descriptors |
| AI contracts and workflows | PhotoAIKit package products | Typed similarity artifacts, Vision/CLIP backends, semantic search, segmentation, and storage contracts |
| Tests | RawCullTests/ and package test targets | Executable behavior contracts, concurrency checks, cache identity checks, and integration coverage |
Architectural Shape
flowchart TD
App["RawCullApp: composition root"] --> Main["RawCullMainView: presentation router"]
Main --> VM["RawCullViewModel: @MainActor orchestration"]
VM --> Feature["Feature models: culling, sharpness, similarity, Deep Review, settings"]
VM --> Actors["Actors: scan, thumbnail, cache, persistence, export"]
Feature --> Analysis["PhotoAnalysisKit"]
Feature --> AI["PhotoAIKit services and typed artifacts"]
Actors --> Parser["RawParserKit"]
VM --> Core["RawCullCore pure models and algorithms"]
Actors --> Storage["Application Support, Caches, selected folders"]The boundaries are based on ownership:
- SwiftUI views render observable state and translate gestures, commands, and
bindings into model actions.
RawCullViewModel and feature view models are @MainActor because their
state drives presentation.- Actors own shared mutable background state and serialize work such as cache
access, thumbnail coalescing, and writes.
RawCullCore contains reusable value models and pure decisions that do not
need UI isolation.RawParserKit, PhotoAnalysisKit, and PhotoAIKit hide specialized parsing
and analysis implementations behind typed boundaries.- Security-scoped access remains alive at the app or operation boundary for as
long as user-selected files are used.
Follow One Feature Through The Code
When investigating a behavior, read it in this direction:
- Find the SwiftUI control or lifecycle trigger.
- Find the
RawCullViewModel or feature-model method it calls. - Identify which actor, adapter, or package owns the expensive or mutable work.
- Follow the value returned to main-actor state.
- Read the tests named after the actor, model, or feature before changing the
invariant.
This approach is more reliable than starting with an individual utility type
because it shows both ownership and the complete lifetime of the operation.
Focused References
1 - Thumbnails and Scan Pipeline
Thumbnails and Scan Pipeline
RawCull separates catalog scanning, thumbnail preloading, UI-driven thumbnail
requests, and full-size preview extraction. They share RawParserKit and
disk-cache infrastructure, but they have deliberately different memory-admission
rules.
The key current invariant is: catalog preloading warms the 200 px grid cache
and disk cache, but only UI demand may admit an image to the preview-size RAM
cache. This keeps scan order from displacing the images the user is actively
viewing. Cache reuse is also replacement-safe: a file written over the same path
is a different source when its size or modification date changes.
Source Map
| Area | Main files |
|---|
| Catalog lifecycle | RawCullViewModel+Catalog.swift, RawCullMainView.swift |
| App/package loading boundary | Model/RawImageLoading.swift, RawParserKit/Sources/RawParserKit/RawImageLoader.swift |
| File discovery and metadata scan | Actors/DiscoverFiles.swift, Actors/ScanFiles.swift |
| Catalog thumbnail preload | Actors/ScanAndCreateThumbnails.swift, Model/Handlers/CreateFileHandlers.swift |
| Preload/UI contention gate | Actors/ThumbnailPreloadGate.swift |
| UI-driven thumbnails | Actors/ThumbnailLoader.swift, Actors/RequestThumbnail.swift, Views/ThumbnailComponents/ThumbnailImageView.swift |
| Thumbnail identity and caches | Model/Cache/ThumbnailCacheKey.swift, Actors/SharedMemoryCache.swift, Actors/DiskCacheManager.swift, Model/Cache/CachedThumbnail.swift |
| Full-size preview loading | Model/FullSizePreviewLoader.swift, Model/Handlers/ZoomPreviewHandler.swift, Actors/FullSizeJPGDiskCache.swift |
| Vendor dispatch | RawParserKit/Sources/RawParserKit/RawFormat.swift, RawFormatRegistry.swift, SonyRawFormat.swift, NikonRawFormat.swift |
| Behavior tests | RawCullTests/ThumbnailLoaderConcurrencyTests.swift, ThumbnailProviderTests.swift, DiskCacheAndScanAdmissionTests.swift, ThumbnailCacheIdentityTests.swift, RawCullVerifyTestsConcurrencyTests.swift |
Catalog Load Flow
flowchart TD
A["User selects catalog folder"] --> B["startCatalogLoad"]
B --> C["Acquire security-scoped access"]
C --> D["ScanFiles.scanFiles"]
D --> E["RawImageLoading.fileMetadata per file"]
E --> F["Sort and filter on MainActor"]
F --> G["Load ratings and valid saved scores"]
G --> H["ScanAndCreateThumbnails.preloadCatalog"]
H --> Gate["ThumbnailPreloadGate: active catalog"]
Gate --> I["200 px grid RAM cache"]
Gate --> J["JPEG thumbnail disk cache"]
F --> K["SwiftUI grid and detail views"]
K --> L["ThumbnailLoader / RequestThumbnail"]
L -. "grid miss waits during preload" .-> Gate
L --> M["Preview RAM cache"]
L --> JCatalog Load Ownership
RawCullViewModel.startCatalogLoad(for:) cancels the older load, starts
security-scoped access, and creates a background-priority catalogLoadTask that
runs handleSourceChange(url:).
handleSourceChange(url:) is main-actor orchestration. File I/O, metadata work,
decoding, and cache I/O are awaited through actors or detached work. After every
significant suspension, isActiveCatalogLoad(_:) and cancellation checks
prevent an old catalog from publishing into the current UI.
Changing catalogs also cancels thumbnail/full-preview work, clears burst
analysis state, stops the old security scope, and releases the current scanning
actors.
File Discovery
DiscoverFiles enumerates a directory off the main actor and filters extensions
through RawFormatRegistry.allExtensions. Its recursive parameter controls
whether subdirectories are traversed.
The current registry contains:
| Format | Extension | Conformer |
|---|
| Sony ARW | .arw | SonyRawFormat |
| Nikon NEF | .nef | NikonRawFormat |
| Adobe DNG | .dng | DNGRawFormat |
ScanFiles performs its own non-recursive directory listing and asks
RawFormatRegistry.format(for:) whether each item is supported. Discovery and
scan therefore share registry policy rather than maintaining separate hard-coded
extension lists.
ScanFiles.scanFiles(url:onProgress:) starts its own security-scoped access and
creates one child task per supported file. Each child:
- reads URL resource values for name, byte size, content type, and modification
date;
- calls the injected
RawImageLoading.fileMetadata(for:) abstraction; - builds a
FileItem with EXIF metadata and the normalized AF point; - records the package focus-location string when available.
The production adapter, RawParserKitImageLoader, maps
RawParserKit.RawImageLoader.metadata(for:) into the app’s ExifMetadata.
RawParserKit reads ImageIO EXIF/TIFF data, dispatches format-specific details
through RawFormatRegistry, and resolves MakerNote or EXIF subject-area focus
evidence. DNG preview discovery classifies TIFF IFDs with NewSubFileType and
Compression before using its legacy positional fallback, so a JPEG-compressed
raw strip is not mistaken for a rendered preview.
This is a single metadata pass per file. The app no longer runs separate EXIF
and MakerNote extraction passes.
If no native focus-location strings are found for the catalog, ScanFiles
falls back to focuspoints.json. This is catalog-wide fallback behavior; it
does not merge JSON entries into a partly populated native result.
FileItem is an app typealias for RawCullCore.RawCullFileItem, keeping scan
output usable by the pure package engines and tests.
Thumbnail Preload
After scan, sort, ratings, and valid saved scores are applied,
ScanAndCreateThumbnails.preloadCatalog(at:targetSize:) walks the catalog again
to warm browsing caches. A catalog URL already recorded in processedURLs is
not preloaded again during the same app session. The actor registers the catalog
with ThumbnailPreloadGate for exactly the lifetime of this preload and always
ends the gate after the task returns or is cancelled.
The outer task group admits at most:
RawImageLoadingConcurrency.thumbnailPreloadLimit
The app-level limit bounds catalog work. RawParserKit adds its own thumbnail
decode limiter and coalesces identical URL/size requests, preventing
simultaneous RAW decodes from growing without bound at the package boundary.
Before any lookup, preload resolves a preview ThumbnailCacheKey from current
file metadata and derives its 200 px grid representation. For each file, preload
then uses this path:
flowchart LR
A["RAW URL"] --> B{"Preview RAM hit?"}
B -->|"yes"| C["Touch preview LRU"]
B -->|"no"| D{"Disk JPEG hit?"}
D -->|"yes"| E["Decode oriented JPEG"]
D -->|"no"| F["RawImageLoading.thumbnailImage"]
C --> G["Populate 200 px grid cache"]
E --> G
F --> G
F --> H["Encode JPEG data"]
H --> I["Background atomic disk save"]All three branches populate the dedicated grid cache when needed. Disk and
source branches do not insert into the preview-size RAM cache.
RequestThumbnail is its only admitter, so preview LRU ordering reflects user
demand instead of scan order. After a cold decode, preload resolves the key
again before caching; if the source changed while decoding, the result is
discarded rather than being stored under stale identity.
On a cold extraction, ScanAndCreateThumbnails converts the actor-owned image
to JPEG Data before launching the background disk save. This avoids sending
CGImage or NSImage across the detached-task boundary.
The Two Memory Caches
SharedMemoryCache owns two NSCache instances:
| Cache | Content | Admission policy |
|---|
Preview cache (memoryCache) | Preview-size images used for detail/list demand | Only RequestThumbnail inserts; a hit touches its LRU position |
Grid cache (gridThumbnailCache) | Downscaled 200 px images | Preload populates it; grid requests can return immediately |
Both caches are cost-limited. The preview item count limit is intentionally high
so byte cost is the normal binding constraint. Under warning memory pressure
both limits shrink to 60%; under critical pressure both caches are cleared and
the preview cache is temporarily capped at 50 MiB.
UI-Driven Thumbnail Loading
ThumbnailImageView has two routes:
.grid calls the shared ThumbnailLoader actor;.list calls RequestThumbnail directly.
For grid targets of 200 px or less, ThumbnailLoader first checks the grid
cache without acquiring a concurrency slot. A miss joins its FIFO-like slot
queue, which allows at most six active requests and removes cancelled waiters
safely.
There is one step before that queue: a grid miss for the catalog currently being
preloaded waits on ThumbnailPreloadGate, then checks the grid cache again.
This prevents the visible grid from launching a duplicate cold decode while
preload is already producing the same catalog. Only grid misses in that catalog
wait; AI indexing, semantic search, Deep Review, model downloads, non-grid
requests, and other catalogs are independent of this gate. Cancellation removes
and resumes the gate waiter safely.
After a slot is acquired, the loader reads saved settings and asks
RequestThumbnail for settings.thumbnailSizePreview. The passed grid target
controls the grid-cache fast path; the slower preview request uses the
configured preview size.
RequestThumbnail resolves a request in this order:
- preview RAM cache;
- oriented JPEG disk cache;
RawImageLoading.thumbnailCGImage source decode.
A disk hit is promoted into preview RAM. A cold source decode is inserted into
preview RAM and asynchronously encoded to the thumbnail disk cache. The cache
also records UI demand, cold extraction, eviction, and “boomerang” diagnostics
for measuring whether a recently evicted image had to be reloaded from disk.
Requests with the same complete ThumbnailCacheKey share one producer in
RequestThumbnail. Each caller has its own continuation. Cancelling one waiter
does not cancel work needed by other waiters; the producer is cancelled only
when its last waiter leaves.
Replacement-Safe Thumbnail Identity
ThumbnailCacheKey is the identity shared by grid RAM, preview RAM, in-flight
request coalescing, and the disk cache. It has two layers:
ThumbnailSourceFingerprint
standardized file URL
file size
modification date
ThumbnailCacheKey
source fingerprint
purpose: grid or preview
requested pixel size
orientation policy
schema version
The source fingerprint is optional. If filesystem size or modification date
cannot be read, RawCull still decodes the image but does not reuse it through a
potentially unsafe key. Purpose and requested size prevent a 200 px grid
representation from satisfying a larger preview request. The schema version
invalidates all older representation semantics without coupling thumbnail
invalidation to PhotoAIKit artifacts.
Disk Thumbnail Cache
DiskCacheManager stores quality-0.7 JPEGs under a directory named for
ThumbnailCacheKey.schemaVersion. The filename is an MD5 hash of the complete,
locale-independent cacheIdentifier: schema, source path bytes, size,
modification-date bits, purpose, requested pixels, and orientation policy. MD5
is used only as a compact filesystem name, not for security.
Older root-level JPEG entries are removed when the current representation-aware
cache is initialized. Loads use OrientationNormalizedImageLoader; a corrupt or
partial JPEG is deleted and treated as a recoverable miss. Writes accept
pre-encoded Data and are atomic, so CGImage does not cross an actor/task
boundary.
Full-Size And Zoom Preview Paths
Full-size inspection does not use the normal thumbnail disk cache.
flowchart TD
A["Zoom / comparison source"] --> B{"Source choice"}
B -->|"Thumbnail"| C["RequestThumbnail"]
B -->|"Embedded JPEG"| D["FullSizePreviewLoader"]
B -->|"Developed RAW"| E["ZoomPreviewHandler"]
D --> F{"Sidecar JPEG exists?"}
F -->|"yes"| G["Load oriented sidecar"]
F -->|"no"| H{"Full-size disk cache hit?"}
H -->|"yes"| I["Load cached embedded preview"]
H -->|"no"| J["RawImageLoading.previewCGImage"]
J --> K["Save embeddedJPG variant"]
E --> L{"developedRAW cache hit?"}
L -->|"no"| M["SonyRawFormat.createFullSizeJPEG"]FullSizePreviewLoader prefers a same-basename .jpg sidecar, then the
full-size disk cache, then RawParserKit extraction. RawParserKit itself
coalesces duplicate preview requests and limits expensive full-size decode work
to two concurrent operations.
The developed-RAW route is currently Sony-specific and uses its own cache
variant. The full-size cache stores quality-0.85 JPEG data and has a versioned
orientation-aware key. Keeping these larger images on disk prevents a few
pixel-peeping previews from evicting many browsing thumbnails from RAM.
App Loading Abstraction
The app depends on the RawImageLoading protocol rather than calling vendor
conformers directly:
| Requirement | Use |
|---|
fileMetadata / exifMetadata | Scan metadata and normalized focus evidence |
thumbnailCGImage / thumbnailImage | On-demand and preload thumbnail generation |
previewCGImage | Embedded/sidecar full-size preview loading |
RawParserKitImageLoader is the production adapter. Tests can inject alternate
loaders without performing real RAW decoding.
Within RawParserKit, RawFormat provides vendor policy:
| Requirement | Use |
|---|
extensions, displayName | ARW/NEF/DNG registry lookup and diagnostics |
extractThumbnail | Vendor thumbnail fallback |
extractEmbeddedPreview | Largest usable embedded preview |
focusLocation | MakerNote AF location |
rawFileTypeString | Compression labels |
sizeClassThresholds, rawSizeClass | Body-specific S/M/L size labels |
extractFullJPEG remains only as a deprecated compatibility requirement; new
code uses extractEmbeddedPreview.
What To Check When Changing This Area
- Preserve the rule that scan/preload does not admit disk or source results to
preview RAM.
- Keep
ThumbnailPreloadGate scoped only to grid misses for the actively
preloading catalog. - Build memory, disk, and coalescing identity from the same complete
ThumbnailCacheKey. - Do not fall back to path-only reuse when source metadata is unavailable.
- Check both the grid-cache fast path and configured preview-size path when
changing thumbnail settings.
- Keep
RawFormatRegistry.allExtensions and RawFormatRegistry.all aligned
when adding a format. - Keep expensive decode limits and cancellation behavior covered by concurrency
tests.
- Convert actor-owned images to
Data before detached cache writes. - Version cache keys when orientation or encoded-image semantics change.
- Run cache-identity, scan-admission, contention, cancellation, and
replacement-at-same-path tests together after changing this pipeline.
- Inspect
isActiveCatalogLoad(_:) guards if stale results appear after
switching folders.
2 - Memory Cache
Cache System
RawCull uses several caches because RAW decoding is expensive and because
different UI surfaces need different representations and lifetimes. The normal
thumbnail path is layered RAM -> disk -> RAW extraction. The zoom path has its
own disk cache for larger embedded JPEGs. Similarity has reusable per-file
artifacts plus a separate catalog-wide burst snapshot. These layers must not
share keys merely because they originate from the same RAW file.
Source Map
| Cache | Main files | Stores |
|---|
| Thumbnail identity | Model/Cache/ThumbnailCacheKey.swift | Source fingerprint plus representation purpose, size, orientation, and schema |
| Preview memory cache | Actors/SharedMemoryCache.swift, Model/Cache/CachedThumbnail.swift | Larger NSImage thumbnails admitted by UI demand |
| Grid memory cache | Actors/SharedMemoryCache.swift | Small 200 px thumbnails populated by preload |
| Thumbnail disk cache | Actors/DiskCacheManager.swift | Representation-aware JPEG thumbnail files under the app cache directory |
| Full-size JPEG disk cache | Actors/FullSizeJPGDiskCache.swift, Model/Handlers/ZoomPreviewHandler.swift | Larger embedded JPEG previews for zoom |
| Per-file similarity artifacts | Actors/PerFileAnalysisArtifactStore.swift | Individually validated Vision/CLIP SimilarityArtifact records |
| Burst analysis cache | Actors/BurstAnalysisCache.swift | Derived catalog snapshots of artifacts, scores, groups, rankings, and review state |
| Cache accounting | Model/Cache/CacheDelegate.swift, Actors/SharedMemoryCache.swift | Lock-backed preview/grid item counts and byte costs kept correct across eviction |
Thumbnail Lookup
flowchart TD
A["Request a purpose and pixel size"] --> Key["Resolve ThumbnailCacheKey"]
Key --> B{"Matching memory representation?"}
B -->|"hit"| C["Return NSImage"]
B -->|"miss"| D{"Matching disk representation?"}
D -->|"hit"| E["Decode JPEG and refill memory"]
D -->|"miss"| F["RawParserKit extractThumbnail"]
F --> G["Create NSImage"]
G --> H["Store preview memory for UI demand"]
G --> I["Store grid memory copy for preload"]
G --> J["Save disk JPEG"]The same lookup order is used by preload and on-demand requests, but admission
differs. ScanAndCreateThumbnails warms disk and the 200 px grid cache after a
catalog opens. ThumbnailLoader and RequestThumbnail resolve UI demand; only
RequestThumbnail admits disk or source results into preview RAM.
ThumbnailPreloadGate prevents a grid miss from duplicating cold work already
being performed by preload for the active catalog.
Cache Identity
ThumbnailCacheKey is the single identity used by memory lookup, disk lookup,
and exact-key request coalescing. It contains:
- a standardized source URL;
- source byte size and modification date;
- representation purpose (
grid or preview); - requested pixel size;
- orientation policy;
- thumbnail schema version.
This answers two separate questions: “Are these still the same source bytes?”
and “Is this the representation the caller requested?” A path-only key fails the
first question when a file is replaced in place. A key without purpose or size
fails the second by allowing a small grid image to satisfy a preview request.
If source metadata cannot be resolved, the key initializer returns nil. The
image may still be decoded for the caller, but RawCull avoids persistent or
coalesced reuse rather than assigning unsafe identity.
ThumbnailCacheIdentityTests covers file replacement, requested-size
separation, purpose separation, orientation policy, and schema participation.
CachedThumbnail
CachedThumbnail wraps an immutable NSImage, its cache cost, and the
standardized source URL associated with the representation.
The cost is calculated from image representations:
cost = sum(rep.pixelsWide * rep.pixelsHigh * 4) * 1.1
The 4 is SharedMemoryCache.costPerPixel, fixed for RGBA. The 1.1
multiplier gives a small overhead buffer for the wrapper and image metadata.
CachedThumbnail is @unchecked Sendable. The project invariant is: construct
the NSImage fully before caching it, then treat it as immutable.
The class intentionally does not conform to NSDiscardableContent. Earlier
versions did, but diagnostics showed NSCache discarded entries too
aggressively and destroyed the RAM hit rate. Eviction is now controlled by
explicit cache limits and memory-pressure handling.
SharedMemoryCache
SharedMemoryCache is an actor singleton:
actor SharedMemoryCache {
nonisolated static let shared = SharedMemoryCache()
}
It is still an actor because it owns configuration, pressure monitoring,
disk-cache references, and statistics. Its two NSCache objects are marked
nonisolated(unsafe) because NSCache is already thread-safe and the app needs
synchronous lookup APIs.
That gives RawCull both:
- actor isolation for mutable app-owned state,
- fast synchronous cache reads for hot thumbnail paths.
Both memory caches are keyed by the complete ThumbnailCacheKey. The preview
and grid APIs remain separate even though they use the same identity type,
making admission policy visible at the call site.
Adaptive Limits
CacheRecommendationPolicy chooses cache caps from physical memory, current
used memory, user maximums, and memory pressure.
Baseline limits:
| Machine memory | Preview baseline | Grid baseline |
|---|
| 64 GB or more | 8000 MB | 2000 MB |
| 32 GB or more | 4096 MB | 1024 MB |
| Less than 32 GB | 2048 MB | 768 MB |
At normal pressure, RawCull calculates available headroom:
freeMB = physicalMB - usedMB
expandableMB = max(0, freeMB - 3072)
extraBudgetMB = expandableMB * 0.5
previewMB = baselinePreview + extraBudgetMB * 0.65
gridMB = baselineGrid + extraBudgetMB * 0.35
Values are rounded up to 256 MB steps and capped by both the machine tier and
user settings.
At warning or critical pressure, the policy falls back toward 60 percent of the
baseline, again respecting user limits and minimum settings.
Eviction Accounting
CacheDelegate is shared by both NSCache instances. When an entry is evicted,
it identifies the originating cache and calls back into SharedMemoryCache to
decrement the corresponding manual cost/count mirrors. This is necessary
because NSCache does not expose current cost or item count. The current tree
does not maintain a separate eviction-history or hit-rate diagnostics store.
Memory Pressure
SharedMemoryCache owns a DispatchSourceMemoryPressure.
| Kernel event | RawCull response |
|---|
.normal | Restore cache configuration from settings/adaptive policy |
.warning | Reduce preview and grid cache cost limits to 60 percent of their current limits |
.critical | Clear both memory caches, reset counters, set preview cache limit to 50 MB until recovery |
The current pressure level is stored behind an OSAllocatedUnfairLock and
exposed as a synchronous nonisolated property. This lets MemoryViewModel
read pressure state without an actor hop.
Disk Caches
DiskCacheManager stores generated quality-0.7 thumbnail JPEGs in a
schema-specific directory. Its filename is an MD5 hash of the complete cache
identifier. MD5 is a fixed-width filesystem key here, not a security primitive.
A corrupt JPEG is removed and treated as a cache miss, and writes are atomic.
Callers encode to Sendable Data before crossing the actor/task boundary.
FullSizeJPGDiskCache stores larger embedded previews for the zoom overlay. It
is deliberately separate from the normal thumbnail cache because zoom images are
larger and should not compete with grid scrolling for memory.
Similarity And Burst Analysis Caches
Similarity persistence has two levels under Application Support:
| Level | Owner | Reuse unit |
|---|
| Per-file artifact store | PerFileAnalysisArtifactStore | One source fingerprint and backend descriptor |
| Burst analysis snapshot | BurstAnalysisCache | One compatible catalog analysis context |
The per-file store lets RawCull reuse valid SimilarityArtifact values when a
catalog-wide snapshot is stale or incomplete. Records include their own schema,
source identity, backend/model descriptor, and pipeline signature. Invalid
records are ignored or removed individually.
BurstAnalysisCache is not an image cache. Its current schema is 9. It stores:
- typed Vision or CLIP similarity artifacts,
- sharpness scores and saliency info,
- burst groups and boundary evidence,
- ranking results,
- review states,
- a digest of the descriptor-and-payload artifact set.
A snapshot is accepted only when the catalog, file count, every file’s
size/modification date, thumbnail size, sharpness descriptor, grouping
configuration, backend and artifact descriptors, artifact schema, pipeline
version, artifact digest, cache schema, and grouping algorithm version still
match. File UUIDs are remapped by path after a new scan. A legacy schema may
supply migration candidates, but it is not accepted as a current snapshot.
What To Check When Changing This Area
- Keep
CachedThumbnail immutable after insertion. - Use
ThumbnailCacheKey consistently for RAM, disk, and in-flight coalescing;
never fall back to path-only reuse. - Keep grid and preview purposes separate and bump the thumbnail schema when
representation semantics change.
- Update manual cost/count mirrors whenever adding or removing cache entries.
- Use Instruments or temporary signposts to distinguish eviction churn from
decoding or I/O regressions; the production tree does not export hit-rate logs.
- Do not feed zoom-preview images into the normal thumbnail RAM cache unless you
intentionally want them competing with grid thumbnails.
- Keep per-file AI artifact persistence separate from the catalog-wide derived
snapshot.
- When changing scoring, similarity, or burst algorithms, update descriptors,
signatures, digests, and schema/algorithm versions so stale analysis data is
rejected.
3 - Memory Pressure
How RawCull measures memory and reacts to macOS pressure events
Memory Pressure
RawCull can hold many images in memory while scanning and culling. The memory-pressure code exists to keep the app responsive when macOS reports pressure and to make cache behavior visible during development.
Source Map
| File | Role |
|---|
Actors/SharedMemoryCache.swift | Owns cache limits, memory-pressure dispatch source, pressure counters, and cache clearing |
Model/ViewModels/MemoryViewModel.swift | Samples total system memory, used system memory, app memory, and pressure threshold for the UI |
Model/Cache/CacheConfig.swift | Contains CacheRecommendationPolicy and cache-limit calculation |
Views/Settings/MemoryTab.swift | Displays memory stats, cache limits, and pressure status |
Views/RawCullSidebarMainView/MemoryWarningLabelView.swift | Shows the sidebar warning during pressure |
Runtime Flow
flowchart TD
A["SharedMemoryCache.ensureReady"] --> B["startMemoryPressureMonitoring"]
B --> C["DispatchSourceMemoryPressure"]
C --> D{"event"}
D -->|"normal"| E["refreshConfig from settings/adaptive policy"]
D -->|"warning"| F["shrink cache limits to 60 percent"]
D -->|"critical"| G["clear memory caches and set 50 MB preview cap"]
F --> H["Notify file handler warning"]
G --> H
E --> I["Clear warning"]SharedMemoryCache starts monitoring lazily during ensureReady(). That is called before thumbnail cache use, so the memory-pressure handler is active before large thumbnail work begins.
Measuring System Memory
MemoryViewModel.getUsedSystemMemory() uses Mach VM statistics:
usedMemory = min((wired + active + compressed) * pageSize, physicalMemory)
The model intentionally uses a simple definition of “used”:
| VM field | Meaning |
|---|
wire_count | Wired pages that cannot be paged out |
active_count | Pages currently active |
compressor_page_count | Pages held by the VM compressor |
The result is clamped to physical memory so the UI cannot show impossible totals.
Measuring App Memory
RawCull measures its own process footprint with:
task_info(... TASK_VM_INFO ...).phys_footprint
That is the value displayed as app memory usage. It is a better signal than simply adding image sizes because it reflects what the kernel says the process currently costs.
Percentages In The UI
MemoryViewModel exposes:
memoryPressurePercentage = memoryPressureThreshold / totalMemory * 100
usedMemoryPercentage = usedMemory / totalMemory * 100
appMemoryPercentage = appMemory / usedMemory * 100
The default pressure threshold is 85 percent of physical memory:
memoryPressureThreshold = totalMemory * 0.85
This UI threshold is informational. The actual pressure response is driven by macOS DispatchSourceMemoryPressure events.
Cache Limit Policy
CacheRecommendationPolicy.adaptiveLimits(...) is the first place to inspect when tuning memory behavior.
RawCull has two independent RAM budgets:
| Cache | NSCache | Primary content | Settings maximum |
|---|
| Preview | memoryCache | Preview and loupe thumbnails | memoryCacheSizeMB |
| Grid | gridThumbnailCache | Grid-size thumbnails | gridCacheSizeMB |
Both use byte cost as the binding constraint. The preview count limit is 10,000 and the grid count limit is 3,000, so normal eviction should be driven by pixel cost rather than item count.
At normal pressure:
- Choose a baseline from physical RAM.
- Estimate free memory from Mach stats.
- Keep a 3 GB reserve.
- Use half of the remaining expandable memory as extra cache budget.
- Split that extra budget 65 percent preview cache and 35 percent grid cache.
- Round to 256 MB steps.
- Clamp by machine tier and user maximums.
The current tiers are:
| Physical RAM | Preview baseline | Grid baseline | Preview tier cap | Grid tier cap | Default user maxima |
|---|
| Less than 32 GB | 2,048 MB | 768 MB | 4,096 MB | 1,024 MB | 4,096 / 1,024 MB |
| 32 GB to less than 64 GB | 4,096 MB | 1,024 MB | 8,000 MB | 2,000 MB | 4,096 / 1,024 MB |
| 64 GB or more | 8,000 MB | 2,000 MB | 8,000 MB | 2,000 MB | 8,000 / 2,000 MB |
The policy never treats a user maximum as a target. It calculates an adaptive recommendation and then clamps preview and grid independently to their saved maxima. At warning or critical state, calculateConfig(from:) uses 60 percent of the tier baseline, rounded up to 256 MB and bounded by the configured minimum and user maximum.
At non-normal pressure, RawCull returns reduced baseline limits and then the live pressure handler may shrink or clear active caches immediately.
Pressure Responses
| Event | Code path | Effect |
|---|
| Normal | handleMemoryPressureEvent, .normal | Set pressure level to normal, increment normal counter, refresh config, clear warning |
| Warning | .warning | Set pressure level to warning, increment warning counter, reduce the current preview and grid cost limits to 60 percent, retain entries for NSCache to evict, show warning |
| Critical | .critical | Set pressure level to critical, increment critical counter, clear preview and grid RAM caches, reset their cost/count mirrors, clear the recent-eviction ring, and set the preview limit to 50 MB |
The grid cache is cleared on critical pressure too. The disk caches are not deleted; only RAM is released.
The 50 MB critical cap applies to the preview cache. Critical handling clears the grid cache but does not rewrite its live cost limit. A later .normal event calls refreshConfig(), re-runs the adaptive calculation using current memory and saved settings, restores both live limits, and clears the UI warning. Recovery does not repopulate either cache; normal demand and preload work warm them again.
Why Pressure State Is Lock-Backed
currentPressureLevel is read by the UI without await. The value is stored in an OSAllocatedUnfairLock, which makes the read/write contract explicit while avoiding a main-actor or cache-actor hop during frequent sampling.
The same pattern is used for the cache cost/count mirrors. The synchronous surface is deliberately narrow: pressure snapshots, NSCache lookups/inserts, and lock-backed cache totals. Configuration, disk-cache access, monitoring setup, and handler ownership remain actor-isolated or explicitly run in detached tasks. This design does not imply that arbitrary cache operations may bypass actor isolation merely because NSCache itself is thread-safe.
Runtime Observability
MemoryTab refreshes MemoryViewModel once per second while the settings view is active. It displays total physical memory, the app’s definition of used system memory, the RawCull process footprint, the informational 85-percent threshold, and the kernel-reported pressure level.
SharedMemoryCache also maintains lock-backed current cost and item-count totals for the preview and grid caches. CacheDelegate decrements those mirrors when NSCache evicts an item, which keeps the values shown in settings accurate. There is no separate memory-diagnostics console or TSV-export pipeline in the current source tree.
Validation When Limits Change
Run the cache-policy tests in RawCullTests/ThumbnailProviderTests.swift. They currently cover the production/testing relationship, explicit CacheConfig values, the 16 GB baseline, expansion and tier caps, user maxima, and warning-state rounding. Add fixtures for every new RAM tier, clamp, rounding rule, or pressure branch.
Also retain the concurrency test in RawCullVerifyTestsDataRaceDetectionTests.swift, which samples currentPressureLevel concurrently.
For an operational limit change, capture a diagnostics session with the same representative catalog and workflow before and after the change:
- Record idle, initial grid population, sustained scrolling, loupe/preview use, and recovery after induced or observed pressure.
- Compare peak process footprint, preview/grid cost and item counts, live limits, and time to first usable grid using Instruments or temporary development instrumentation.
- Confirm warning shrinks both live caches without deleting disk data.
- Confirm critical clears both RAM caches, preview falls to the 50 MB cap, counters record the event, and
.normal restores adaptive limits. - Reject a larger limit if it only raises footprint without improving reuse; reject a smaller limit if it creates repeated cold extraction or visible grid/preview churn.
What To Check When Changing This Area
- Memory sampling should stay off the main actor; Mach calls run in
Task.detached. - Do not rely only on the 85 percent UI threshold; the real emergency signal is the kernel pressure event.
- If cache limits look strange, inspect both user settings and the adaptive tier caps.
- Treat preview and grid limits separately; a healthy total can hide churn in one cache.
- The settings display samples state; use Instruments or signposts when a change needs event-level evidence.
- Critical pressure should free RAM quickly and should not delete disk caches.
- After recovery, verify both live limits were recalculated from current settings and memory, not merely reset to hard-coded values.
4 - Concurrency
Concurrency
This page describes the current RawCull runtime in the RawCull repository.
The app uses Swift 6 with default main-actor isolation: presentation state
begins on @MainActor, shared mutable background state lives in actors, and
CPU, decode, or file work leaves inherited UI isolation explicitly.
A Swift task is not a thread. await is a suspension point, not an automatic
hop to a background thread. RawCull reasons about actor, executor, queue, and
task ownership; it never uses worker-thread identity as an application
invariant.
Composition And Presentation Boundaries
RawCullApp retains the two roots returned by RawCullApplicationState.live():
one RawCullViewModel and one RawCullIntelligenceRuntime. The application
state assembles one integration, one shared similarity model, focused similarity
and semantic-search features, a Deep Review controller,
settings/model-management models, and the main view model. Settings publishes a
complete revisioned configuration to the runtime rather than invoking
independent callbacks on the view model.
RawCullMainView is the presentation boundary. It binds the main-actor view
model to SwiftUI scenes, sheets, alerts, commands, and child views. It does not
own catalog decoding, thumbnail admission, analysis, or persistence work.
RawCullViewModel is @Observable @MainActor. It owns the current catalog,
selection, application operation lifetimes, result application, culling policy,
and active catalog security scope. BurstAnalysisCoordinator owns the burst
worker task, generation, progress, cache preparation, compute orchestration, and
cache save. RawCullSimilarityFeature, RawCullSemanticSearchFeature, and
DeepAIReviewController are the view-facing AI boundaries. Background actors
own mutable file, cache, provider, and persistence state.
flowchart LR
APP["RawCullApp<br/>stable @State roots"] --> STATE["RawCullApplicationState<br/>object-graph assembly"]
STATE --> VM["@MainActor RawCullViewModel<br/>catalog and application policy"]
STATE --> RT["@MainActor RawCullIntelligenceRuntime<br/>AI lifetime and configuration"]
RT --> AIS["RawCullAISettingsModel"]
RT --> FEATURES["Focused similarity, semantic, and Deep Review surfaces"]
VM --> VIEW["RawCullMainView and SwiftUI presentation"]
VM --> BURST["BurstAnalysisCoordinator"]
FEATURES --> ACTORS["Background actors and package services"]
BURST --> ACTORS
ACTORS --> VALUES["Sendable results"]
VALUES --> FEATURES
VALUES --> BURSTFour Runtime Paths And Their Hops
flowchart TB
subgraph CATALOG["Catalog load"]
C1["SwiftUI selection<br/>MainActor"] --> C2["flushPersistence<br/>CullingModel"]
C2 --> C3["start catalog security scope<br/>RawCullViewModel"]
C3 --> C4["catalogLoadTask"]
C4 -->|"await actor"| C5["ScanFiles actor"]
C5 -->|"bounded task group"| C6["metadata and focus-point reads"]
C6 --> C7["MainActor publication<br/>only if catalog is still active"]
end
subgraph THUMBS["Visible thumbnail demand"]
T1["SwiftUI task"] -->|"await actor"| T2["ThumbnailLoader"]
T2 -->|"grid miss only"| T3["ThumbnailPreloadGate"]
T2 --> T4["six-slot admission"]
T4 --> T5["RequestThumbnail exact-key single flight"]
T5 --> T6["RAM / disk / RawParserKit decode"]
T6 --> T7["CGImage result; caller rechecks cancellation"]
end
subgraph BURST["Burst indexing"]
B1["RawCullViewModel builds immutable request"] --> B2["BurstAnalysisCoordinator generation + task"]
B2 --> B3["cache prepare / bounded scoring and indexing"]
B3 --> B4["provider and artifact-store actors"]
B4 --> B5["grouping and ranking"]
B5 --> B6["callback publishes only if generation + catalog match"]
B6 --> B7["repository-backed cache commit"]
end
subgraph TERM["Application termination"]
A1["AppDelegate.applicationShouldTerminate<br/>MainActor"] --> A2["terminateLater"]
A2 --> A3["terminationTask"]
A3 --> A4["await CullingModel.flushPersistence"]
A4 -->|"success"| A5["stop active catalog scope"]
A4 -->|"failure"| A6["reply false; keep app and scope alive"]
A5 --> A7["reply true"]
end
C7 ~~~ T1
T7 ~~~ B1
B7 ~~~ A1The invisible ordering links (~~~) make Mermaid place the four runtime paths
below one another. They describe layout only; they are not runtime hops.
The catalog path is latest-wins. startCatalogLoad first flushes pending
culling data, then checks cancellation and the selected source before starting
the new load. cancelCatalogLoad cancels catalog, preload, hydration, and
cache-warming tasks; asks batch actors to cancel their inner tasks; resets
dependent analysis; and stops the active catalog scope. Publication repeatedly
checks isActiveCatalogLoad(url) and task cancellation.
The termination path deliberately delays AppKit termination. AppDelegate
stores one termination task, flushes persistence, and releases the catalog
security scope only after a successful save. A failed flush returns false to
AppKit, leaving the app open so the user can retry without losing access or
unsaved state.
Concurrency Mechanisms
Main-Actor Isolation
Use the main actor for observable UI state and workflow coordination, not for
expensive work.
| Owner | State and lifetime owned |
|---|
RawCullViewModel | Active catalog, selection, security scope, catalog/preload/export tasks, result application, culling and presentation policy |
CullingModel | In-memory culling records, debounced save task, persistence revision and errors |
SharpnessScoringModel | Scoring request/generation, options, progress, scores and breakdowns |
SimilarityScoringModel | Artifact, search, ranking, and grouping generations and published results |
RawCullSimilarityFeature / RawCullSemanticSearchFeature | Focused operation surfaces, application bindings, and projections over the one shared scoring model |
BurstAnalysisCoordinator | Burst worker task, generation, progress, cache preparation, missing compute, grouping, ranking, and cache save |
DeepAIReviewController / DeepAIReviewFeature | App-facing request adaptation plus review generation, progress, recommendation, and cancellation |
RawCullIntelligenceRuntime | Stable AI feature lifetimes and last accepted configuration revision/identity |
RawCullAISettingsModel | Capability-refresh generation, preferences, and complete configuration publication |
A task created from one of these owners inherits main-actor isolation. That is
useful for its stateful prefix and final publication. Heavy work must then be
reached through an actor, a package async API, or an explicit @concurrent
boundary.
Actor Serialization
Actors protect shared asynchronous state.
| Actor | Serialized invariant |
|---|
ThumbnailPreloadGate | Active preload catalogs and parked grid-demand waiters |
ThumbnailLoader | Six decode slots, FIFO continuations, and cached settings |
RequestThumbnail | Setup single flight and one producer per exact ThumbnailCacheKey |
ScanFiles, ScanAndCreateThumbnails, ExtractAndSaveJPGs | Batch task, progress, cancellation, and result accumulation |
DiskCacheManager, BurstAnalysisCache, WriteSavedFilesJSON | File-backed cache or persistence transactions |
| PhotoAIKit provider/store actors | Model runtime, artifact, index, and mask-store state |
Actor serialization does not make independent work parallel. Where file work can
overlap, owners use a bounded task group and keep only a sliding window of
children active.
Bounded Structured Parallelism
| Operation | Bound |
|---|
| Visible thumbnail loading | 6 slots in ThumbnailLoader |
| App scan/extract batches | RawImageLoadingConcurrency.batchExtractionLimit |
| Sharpness scoring | Fast 6, Balanced 4, High Precision 3; RAW demosaic capped at 2 |
| PhotoAnalysisKit batch and calibration | Caller supplied; RawCull passes the scoring bound |
| PhotoAIKit indexing | maximumConcurrentTasks sliding window |
Structured child tasks inherit cancellation and finish before the task-group
scope returns. Completion order can drive progress, but final arrays are
restored to request order where order is part of the API. Cancellation calls
cancelAll() and partial results are not committed as a successful run.
Explicit Concurrent Work
A task marked @concurrent, or an @concurrent function, explicitly leaves
inherited actor isolation while retaining structured task behavior. Examples
include RawCull’s RAW demosaic helper, sorting and ranking helpers, diagnostics,
and PhotoAnalysisKit’s Core Image/Vision worker. The pinned PhotoAnalysisKit
1.3.1 implementation uses an explicit concurrent child and a cancellation
handler; it does not use a detached task for the focus engine.
Detached And Queue-Backed Work
Task.detached is reserved for operations that deliberately must not inherit
actor isolation, including selected cache/file operations and conversion of
blocking image representations. The owner still awaits or otherwise controls the
result, and propagates cancellation where required.
Blocking ImageIO work is stronger than ordinary CPU work. RawParserKit bridges
it to a GCD queue with a checked continuation and lock-backed cancellation state
so it does not occupy Swift’s cooperative executor. Framework pipe and PTY
callbacks similarly arrive outside main-actor isolation and explicitly return
through a main-actor task.
Thumbnail Contention Invariants
Thumbnail concurrency has three separate layers:
ThumbnailPreloadGate blocks only grid misses whose parent directory is the
catalog currently being preloaded. Other catalogs, preview requests, AI
workflows, semantic search, and Deep Review bypass the gate.ThumbnailLoader admits at most six active loads. A released slot is
transferred directly to the next waiter; cancellation removes and resumes a
queued continuation exactly once.RequestThumbnail coalesces only an exact ThumbnailCacheKey. Each waiter
owns its continuation. Cancelling one waiter preserves the shared producer;
cancelling the final waiter removes the table entry and cancels the producer.
ThumbnailCacheKey is replacement-safe and representation-aware. It includes
the standardized path, source file size, modification-date bit pattern, purpose
(grid or preview), requested pixel size, orientation policy, and thumbnail
schema. If source metadata cannot be resolved, reusable coalescing and caching
are skipped: a path alone cannot identify bytes replaced in place.
Every in-flight request also has a generation UUID. A producer may finish after
its final waiter was cancelled and a new request for the same key began; the
generation check prevents that old completion from removing or resuming the new
request’s waiters.
Latest-Wins And Stale-Result Prevention
Cancellation is cooperative and is not enough by itself. A callback or framework
operation may complete after cancellation, so owners pair task cancellation with
identity checks.
| Workflow | Commit guard |
|---|
| Catalog load | Active catalog URL, selected source, and cancellation |
| Catalog sort/search | Requested text, sort order, file IDs, and cancellation |
| Sharpness scoring | ScoringRequest, scoring-generation UUID, and cancellation |
| Similarity hydrate/index/rank/search/group | Per-workflow generation plus catalog/query snapshots |
| Burst analysis | Burst generation plus catalog URL |
| Deep Review | Integer generation plus cancellation |
| Raw diagnostics | Generation UUID plus cancellation |
| Thumbnail single flight | Exact cache key plus producer-generation UUID |
Progress callbacks use the same guard as final publication. An older run must
not update a newer run’s progress, clear its task property, or publish a stale
result.
Continuation Ownership And Cleanup
| Owner | Normal resume | Cancellation cleanup |
|---|
ThumbnailPreloadGate | The matching catalog preload ends | Remove waiter, record cancellation, resume false |
ThumbnailLoader | A real slot is transferred | Remove queued waiter and resume cancelled |
RequestThumbnail | Exact-key producer finishes | Remove waiter and resume nil; cancel producer when none remain |
| RawParserKit decode limiter | A decode permit becomes free | Remove cancelled admission waiter and resume cancellation |
| RawParserKit ImageIO bridge | GCD operation completes | Lock-backed state guarantees exactly one completion |
The collection owner is responsible for exactly one continuation resume. Never
hold a lock across await, invoke arbitrary client code while holding a lock,
or use a continuation without a cancellation path.
Security-Scope And Persistence Lifetimes
| Operation | Owner | Start | Stop / failure cleanup |
|---|
| Active catalog | RawCullViewModel | Before catalogLoadTask, after the previous catalog is flushed | Catalog cancellation, empty/failed load, replacement, successful termination, or deinit |
| Selected JPG export | RawCullViewModel+Thumbnails | Before constructing ExtractAndSaveJPGs | Main-actor completion path after success, failure, or cancellation |
| Rsync copy | ExecuteCopyFiles | Active catalog URL for source; destBookmark for destination | One idempotent cleanup() on startup failure, completion, close, or deinit; it also removes the operation’s include-list file |
| App termination | AppDelegate and CullingModel | No new scope; uses the active catalog scope | Flush first; release scope only on successful flush |
Every successful security-scope start has one owner that records the URL and one
idempotent cleanup path.
Source-To-Test Map
| Runtime rule | Protecting tests |
|---|
| Exact-key thumbnail coalescing and independent waiter cancellation | RawCullTests/ThumbnailProviderTests.swift |
| Preload-gate drain and cancellation races | RawCullTests/RawCullVerifyTestsConcurrencyTests.swift |
| Six-slot limit, FIFO transfer, queued cancellation, and cancel-all | RawCullTests/ThumbnailLoaderConcurrencyTests.swift |
| Replacement-safe source/representation identity | RawCullTests/ThumbnailCacheIdentityTests.swift |
| Scoring request coalescing and generation-safe completion | RawCullTests/SharpnessScoringTests.swift |
| Persistence debounce, failed-save dirty state, retry, atomic backup/write, and legacy decode | RawCullTests/CullingModelTests.swift |
| Catalog scope idempotence, replacement, failed start, cancellation, and empty catalog | RawCullTests/RawCullVerifyViewModelSecurityScopeTests.swift |
| Export selection and denied destination scope | RawCullTests/ExtractJPGsSelectionTests.swift |
| Rsync source/destination cleanup and operation-local include files | RawCullTests/ExecuteCopyFilesStartupTests.swift |
| Package batch ordering, bounds, progress, and cancellation | PhotoAnalysisKit/Tests/PhotoAnalysisKitTests/PhotoAnalysisBatchTests.swift |
Review Checklist
- Which actor or task owns the mutable state and lifetime?
- Does the synchronous prefix require that actor?
- Is independent work bounded and structured?
- Is the work CPU-bound, blocking, or callback-based, and which executor or
queue boundary is appropriate?
- What Sendable value crosses the boundary?
- How is every continuation resumed on success and cancellation?
- Which generation, catalog, query, key, or request identity prevents a stale
commit?
- Who stops each successfully started security scope?
- What test makes the invariant executable?
“await moves this to a background thread” is not a valid concurrency model.
Name the actual actor, task, executor, queue, cancellation owner, and commit
guard.
5 - Focus Mask and Sharpness
Focus Mask And Sharpness
RawCull exposes four related results, but they are not interchangeable:
| Result | Meaning | Main consumer |
|---|
| Scalar sharpness | A package-computed ranking value based on full-frame, salient-subject, AF, and local-patch detail | Sorting, burst ranking, persistence |
| Saliency evidence | Vision candidates and an optional classification label/confidence | Subject selection and diagnostics |
| Focus-point evidence | Camera AF position plus AF-center, neighborhood, and local scores | Region selection and diagnostics |
| Rendered focus mask | A thresholded, colorized overlay clipped to selected evidence patches | Visual inspection only |
The scalar score and the overlay share edge-energy and evidence machinery in
PhotoAnalysisKit, but showing more red pixels does not increase a stored score.
Mask presentation does not change scalar analysis. A weak or unfocused image is
allowed to produce an empty overlay; the renderer does not lower its threshold
merely to manufacture visible evidence.
RawCull pins PhotoAnalysisKit 1.3.1, revision
2a1466e04d821fa2628d6985296643e0d0c7e465. The package owns sharpness,
saliency, calibration, focus evidence, and mask algorithms. RawCull owns UI
settings, file decoding/source selection, workflow lifetime, normalization, and
persistence.
Ownership From Controls To Package
flowchart TD
CONTROLS["SharpnessControlsView<br/>start, cancel, photo type"] --> VM["@MainActor RawCullViewModel<br/>target files and persistence"]
SHEET["ScoringParametersSheetView<br/>quality, source, size, config"] --> MODEL["@MainActor SharpnessScoringModel<br/>request, generation, progress, results"]
SETTINGS["FocusSettingsTab and FocusMaskControlsView"] --> FM["@MainActor FocusMaskModel<br/>presentation bridge"]
VM --> MODEL
MODEL --> ADAPTER["RawCullPhotoAnalysisAdapter<br/>host file/source adapter"]
ADAPTER --> LOAD{"RawCull-owned input source"}
LOAD -->|"embeddedPreview"| RPK["RawParserKitImageLoader"]
LOAD -->|"rawDemosaic"| RAW["CIRAWFilter concurrent worker"]
RPK --> INPUT["PhotoAnalysisInput<br/>CGImage + ISO + aperture + AF"]
RAW --> INPUT
INPUT --> ANALYZER["PhotoAnalysisKit.PhotoAnalyzer"]
ANALYZER --> RESULTS["PhotoAnalysisResult / batch result<br/>score, saliency, breakdown"]
ANALYZER --> MASK["CGImage focus mask"]
RESULTS --> MODEL
MASK --> FM
MODEL --> PERSIST["CullingModel and savedfiles.json"]SharpnessScoringModel and FocusMaskModel are @Observable @MainActor
because they publish UI state. Their PhotoAnalyzer and adapter values are
immutable, nonisolated boundaries. PhotoAnalysisKit’s internal FocusMaskEngine
is @unchecked Sendable only because Core Image does not declare CIContext
Sendable; the documented invariant is an immutable engine with value snapshots
for every operation.
Configuration Resolution
A scoring run snapshots one effective SharpnessConfiguration:
focusMaskModel.config
-> SharpnessPhotoType.packagePreset.applying
-> SharpnessScoringQuality.packageQuality.applying
-> PhotoAnalyzer applies each input's ISO and aperture hint
RawCull’s default focus configuration is .birdsInFlight. Photo type maps to
PhotoAnalysisKit presets: Automatic, Birds and Wildlife, Portrait, Landscape, or
General Action. Quality maps to Fast, Balanced, or High Precision. Automatic
preserves the
shared config rather than selecting a preset from image classification; the
default .birdsInFlight has no explicit weight override, unlike the explicit
Birds and Wildlife preset.
The selected thumbnail setting is normalized before decode:
effective size = min(max(user value, quality minimum), 2048)
quality minimum:
Fast 512
Balanced 768
High Precision 1024
The UI currently offers 1024, 1536, and 2048 px choices. The quality minimum
still matters for old settings and programmatic values. Embedded previews use
RawParserKit’s thumbnail loader. RAW demosaic uses CIRAWFilter, disables added
sharpness, and sets detail 0.6, contrast 1.0, and exposure 0 before scaling the
longest side to the effective size.
Analysis Descriptor And Cache Identity
PhotoAnalyzer.sharpnessDescriptor(for:) produces
SharpnessAnalysisDescriptor, the package-owned identity for non-mask scalar
analysis. At the pinned PhotoAnalysisKit 1.3.1 revision it has:
- descriptor schema version 1;
- scalar algorithm version 4;
- ISO-scaling policy version 1;
- aperture-hint policy version 1;
- the scoring-affecting configuration values and the stable scoring gain 7.62.
It deliberately excludes per-image ISO and aperture, decoded image size, input
source, source-file identity, and mask-only presentation settings. RawCull adds
the missing host identity in SharpnessScoringSignature:
| Identity layer | Fields |
|---|
| Package descriptor | Algorithm/policy versions, scoring configuration, stable gain |
| RawCull signature | Descriptor + embedded-preview/RAW source + effective maximum pixel size |
| Per-file persistence validation | Signature + source file size + modification date |
Legacy signatures without a package descriptor still decode, but compare stale
to every current signature. On catalog load, RawCull restores a score only when
the full signature matches and the current file size and modification date match
(date tolerance is 0.001 seconds). A change to scalar-affecting config, quality,
source, size, algorithm identity,
or source-file metadata therefore forces recomputation.
Calibration Lifetime
Calibration is a visual-threshold operation, not catalog normalization of
the scalar score. Before scoring, RawCull asks FocusMaskModel to load inputs
and call PhotoAnalyzer.calibrate, using a dedicated 1616 px maximum rather
than the scoring thumbnail size. PhotoAnalysisKit:
- loads inputs with the same bounded concurrency and source choice as scoring;
- applies each file’s ISO and aperture;
- disables classification;
- samples up to about 4096 positive Laplacian energies per successful image;
- requires at least five successful images;
- chooses the requested percentile (RawCull uses 0.90) and clamps the threshold
to 0.01…0.95.
Only focusMaskModel.config.threshold is updated. Scalar scoring uses the
stable gain and package descriptor; it does not depend on the catalog’s
calibration distribution. The calibrated threshold remains in the shared focus
model until settings, later calibration, or model reset changes it.
Scoring, Progress, Cancellation, And Publication
SharpnessScoringModel.scoreFiles builds a request identity from ordered file
IDs, the scoring signature, and the concurrency limit.
- An identical request already in flight is coalesced and awaited.
- A different request cancels the old task and installs a new generation UUID.
- Fast, Balanced, and High Precision admit 6, 4, and 3 package tasks
respectively; RAW demosaic is capped at 2.
- PhotoAnalysisKit keeps a sliding task-group window, reports completion-order
progress, and restores final results to request order.
- RawCull publishes progress only while the generation matches.
- Cancellation makes the package return nil and partial results are discarded.
- Final score, saliency, and breakdown dictionaries are replaced only if the
task is not cancelled and the generation still matches.
- A successful run enables sharpness sorting, then
RawCullViewModel merges the
results into CullingModel.
This is latest-wins state management. Cancelling work alone is insufficient;
every progress and final commit also checks the generation.
Mask tasks have the same presentation rule. SwiftUI views own stored mask tasks,
cancel them on image/config replacement, and check cancellation before assigning
the returned overlay or diagnostics.
Raw Score Versus UI Label
PhotoAnalysisKit does not promise that the raw scalar is a percentage or clamp
the final value to 0…1. It is a relative detail metric whose scale is kept
stable by the package gain and descriptor.
RawCull computes an O(1) UI denominator when the score dictionary changes:
fewer than 2 scores -> lone score, or 1
2 through 9 -> maximum
10 or more -> element at floor((n - 1) * 0.90) in sorted scores
denominator floor -> 1e-6
Badge consumers clamp score / maxScore to 0…1 before mapping it to
presentation labels. A lone score above the denominator floor therefore displays as 100%, and
an all-soft catalog can still produce Sharp labels. The labels are relative to the
current score set and do not establish absolute focus quality. That UI
normalization is not persisted as the package score.
Focus Mask Presentation
FocusMaskModel offers two package-backed paths:
| Path | Package call | Use |
|---|
| Existing image, optional saved evidence | PhotoAnalyzer.focusMask | Render without repeating classification and, when evidence includes a winning saliency rectangle, without repeating saliency selection |
| Image plus fresh diagnostics | PhotoAnalyzer.analyzeWithFocusMask | Compute scalar/saliency/evidence and render one aligned mask |
Views pass a CGImage, ISO, aperture, normalized AF point, scale, and a
configuration snapshot. RawCull adapts the package breakdown only to add
SharpnessScoringSource; it does not reinterpret the numeric fields.
Presentation-only controls include threshold, dilation, erosion, feathering,
raw-Laplacian display, and subject isolation. The legacy
guaranteeVisibleFocusEvidence and minimumEvidenceCoverage properties remain
source-compatible but no longer relax rendering. Some shared values such as pre-blur, border inset,
AF radii, ISO, and aperture affect the evidence image or regions used by both
paths. See Detailed Focus Mask Computation for the exact
stage classification.
Read These Files In Order
- RawCull/Views/GridView/SharpnessControlsView.swift and
ScoringParametersSheetView.swift — user entry and settings.
- RawCull/Model/ViewModels/RawCullViewModel+Sharpness.swift — target scope,
persistence handoff, and sort.
- RawCull/Model/ViewModels/FocusandSharpness/SharpnessScoringModel.swift —
request identity, generation, bounds, progress, and publication.
- RawCull/Model/ViewModels/FocusandSharpness/SharpnessScoringOptions.swift
— RawCull-to-package preset, quality, source, and size mapping.
- RawCull/Model/ViewModels/FocusandSharpness/RawCullPhotoAnalysisAdapter.swift
— decode ownership and
PhotoAnalysisInput construction. - RawCull/Model/ViewModels/FocusandSharpness/FocusMaskModel.swift and
FocusMaskTypes.swift — mask bridge and result adaptation.
- PhotoAnalysisKit/Sources/PhotoAnalysisKit/PhotoAnalyzer.swift and
PhotoAnalysisBatch.swift — public package facade and bounded batch.
- SharpnessConfiguration.swift, SharpnessPresets.swift, and
SharpnessAnalysisDescriptor.swift — defaults, policy, and identity.
- FocusMaskEngine+Scoring.swift, FocusMaskEngine+MaskGeneration.swift,
and Resources/Kernels.ci.metal — algorithm implementation.
Continue with Detailed Sharpness Scoring for the
numeric algorithm, or Detailed Focus Mask Computation
for overlay rendering.
Protecting Tests
| Boundary or invariant | Tests |
|---|
| RawCull/package Metal integration and scoring-source adaptation | RawCullTests/PhotoAnalysisKitIntegrationTests.swift |
| Preset/quality mapping, descriptor signatures, legacy invalidation, coalesced runs, UI normalization | RawCullTests/SharpnessScoringTests.swift |
| Persistence merge and file/signature validation behavior | RawCullTests/CullingModelTests.swift |
| Batch bounds, ordering, progress, decode failure, cancellation | PhotoAnalysisKitTests/PhotoAnalysisBatchTests.swift |
| Descriptor inclusions/exclusions and policy versions | PhotoAnalysisKitTests/SharpnessAnalysisDescriptorTests.swift |
| Numeric score helpers, aperture policy, ISO curve, focus-failure classification | PhotoAnalysisKitTests/SharpnessMetricsTests.swift |
| Analyze, mask, calibration, and cancellation facade | PhotoAnalysisKitTests/PhotoAnalyzerTests.swift |
6 - Detailed Sharpness Scoring
Detailed Sharpness Scoring
This is the algorithm-level reference for scalar sharpness at PhotoAnalysisKit
1.3.1, revision 2a1466e04d821fa2628d6985296643e0d0c7e465. RawCull chooses
files and image sources, maps UI settings, bounds work, publishes progress,
normalizes badges, and persists results. PhotoAnalysisKit owns the analysis
formula.
For ownership, cancellation, calibration, and UI flow, start with
Focus Mask And Sharpness. This page concentrates on the package
algorithm and names RawCull policy only where it changes package input or
consumes package output.
Compact Worked Example
Consider a wildlife frame after Laplacian sampling:
full-frame robust score 0.18
AF broad score 0.30
Vision salient broad score 0.24
best AF-local patch 0.34
best salient-interior patch 0.28
wildlife salient weight 0.85
subject saliency area ignored because AF exists
subject border fraction 0.50 (below silhouette trigger)
subject micro-contrast 0.020
aperture f/5.6 -> wide gate 0.010...0.025
The actual pipeline combines the evidence in stages:
broad subject = 0.6 * 0.30 + 0.4 * 0.24 = 0.276
local detail = 0.6 * 0.34 + 0.4 * 0.28 = 0.316
subject = 0.75 * 0.276 + 0.25 * 0.316 = 0.286
base = 0.15 * 0.18 + 0.85 * 0.286 = 0.2701
blur t = (0.020 - 0.010) / (0.025 - 0.010) = 0.6667
attenuation = 0.20 + 0.80 * 0.6667 = 0.7333
final = 0.2701 * 0.7333 = about 0.198
There is no silhouette penalty because 50% is below 62%, and no subject-size
bonus because an AF region exists. The final value is a raw relative score, not
a percentage. RawCull later normalizes it against the catalog denominator for a
badge.
The numeric values above are illustrative inputs to the current formulas; they
are not fixture output from a particular image.
End-To-End Boundary
flowchart TD
UI["RawCull controls"] --> RVM["RawCullViewModel target files"]
RVM --> SM["SharpnessScoringModel<br/>preset + quality + source + size"]
SM --> ADAPTER["RawCullPhotoAnalysisAdapter"]
ADAPTER --> DECODE{"Host-owned decode"}
DECODE -->|"embedded"| RPK["RawParserKitImageLoader"]
DECODE -->|"RAW"| CIRAW["CIRAWFilter"]
RPK --> INPUT["PhotoAnalysisInput"]
CIRAW --> INPUT
INPUT --> BATCH["PhotoAnalyzer.analyzeBatch"]
BATCH --> ANALYZE["PhotoAnalyzer.analyze"]
ANALYZE --> VISION["Vision saliency/classification"]
ANALYZE --> LAPLACIAN["Metal Laplacian + robust region metrics"]
VISION --> BLEND["subject/full blend, adjustments, blur gate"]
LAPLACIAN --> BLEND
BLEND --> RESULT["PhotoAnalysisResult / SharpnessBreakdown"]
RESULT --> SM
SM --> PERSIST["RawCull normalization, sorting, persistence"]Target Files
RawCullViewModel.sharpnessScoringTargetFiles chooses, in order:
- explicitly selected files, keeping visible selected files first and appending
selected files hidden by the current projection;
- the exact active star-rating set;
- otherwise the active catalog set, including semantic-search scope, sorted by
filename.
Calibration and scoring use the same target array and image source. Calibration
uses a separate maximum size of 1616 px, while scoring uses the selected effective
size. Calibration adjusts only the visual mask threshold; it does not calibrate
the scalar score. After a non-cancelled,
nonempty result, RawCull merges scores and subject labels into CullingModel
and reapplies catalog sorting.
Protected by: RawCullTests/CullingModelTests.swift target-scope and
calibrateAndScoreCurrentCatalog tests.
Preset And Quality
RawCull starts from the shared focus config, initially .birdsInFlight, then
applies the selected PhotoAnalysisKit preset and quality. Auto preserves that
config; it does not infer a photo type from classification. Unlike the explicit
Birds/Wildlife preset, .birdsInFlight leaves the explicit salient override nil,
so the f/8 aperture fallback can reduce Auto’s subject weight to 0.55.
| RawCull photo type | Package changes from the input config |
|---|
| Auto | No preset change |
| Birds/Wildlife | pre-blur 2.2; border 0.05; salient weight and explicit override 0.85; size factor 0.05; silhouette strength 0.55; AF radius 0.06; classify and isolate |
| Portrait | pre-blur at most 1.7; salient weight/override 0.80; size factor 0.08; silhouette 0.25; AF radius 0.10; classify and isolate |
| Landscape | pre-blur at most 1.55; salient weight 0.35 with explicit override 0.35; size factor 0; silhouette 0.15; AF radius 0; do not isolate mask |
| Action | pre-blur 2.0; salient weight/override 0.65; size factor 0.05; silhouette 0.40; AF radius 0.09; classify and isolate |
| Quality | Package fine-detail weight | RawCull minimum size | RawCull concurrency |
|---|
| Fast | 0 | 512 | 6 |
| Balanced | at least 0.25 | 768 | 4 |
| High Precision | at least 0.45 and classification enabled | 1024 | 3 |
The effective maximum pixel size is clamped between the quality minimum
and 2048; a nonpositive setting resolves to 2048. RAW demosaic further caps
concurrency at 2.
Protected by: RawCullTests/SharpnessScoringTests.swift preset, quality, size,
and concurrency assertions; package policy also has
PhotoAnalysisKitTests/SharpnessMetricsTests.swift.
Image Source
| Source | RawCull implementation |
|---|
| Embedded Preview | RawParserKitImageLoader.shared.thumbnailCGImage at the effective size |
| RAW Demosaic | Concurrent CIRAWFilter decode; sharpness 0, detail 0.6, contrast 1.0, exposure 0; scale longest side to effective size |
The adapter returns PhotoAnalysisInput(image:iso:aperture:normalizedAFPoint:).
PhotoAnalysisKit never opens RawCull files and does not choose a source.
Protected by: RawCullTests/PhotoAnalysisKitIntegrationTests.swift and adapter
overrides in RawCullTests/SharpnessScoringTests.swift.
2. Package Defaults And Per-Image Resolution
The default SharpnessConfiguration values at the pinned revision are:
| Value | Default | Role |
|---|
preBlurRadius | 1.92 | Shared edge-energy input |
iso | 400 | Replaced from each input |
threshold | 0.46 | Mask presentation only |
energyMultiplier | 7.62 | Stable scoring and shared edge gain |
borderInsetFraction | 0.04 | Scalar full-frame border exclusion; also blackens mask border |
salientWeight | 0.75 | Scalar subject/full blend |
subjectSizeFactor | 0.10 | Scalar saliency-only size bonus |
silhouettePenaltyStrength | 0.55 | Scalar penalty maximum |
fineDetailBlendWeight | 0 | Scalar second pass |
enableSubjectClassification | true | Label only; saliency candidates are still analysis evidence |
afRegionRadius | 0.12 | Broad AF scoring region |
afCenterRegionRadius | 0.025 | Evidence diagnostic/local region |
afNeighborhoodRegionRadius | 0.075 | Evidence diagnostic/local-patch region |
apertureHint | mid | Replaced from each input |
PhotoAnalyzer.analyze copies the config, sets the input ISO, and derives:
| Aperture | Hint | Blur gate low/high | Blur damp | Fallback salient override |
|---|
| f/5.6 or wider | wide | 0.010 / 0.025 | 1.0 | none |
| between f/5.6 and f/8 | mid | 0.008 / 0.022 | 1.0 | none |
| f/8 or narrower | landscape | 0.006 / 0.018 | 0.8 | 0.55 |
| missing | mid | 0.008 / 0.022 | 1.0 | none |
An explicit preset override wins over the aperture fallback.
Protected by: PhotoAnalysisKitTests/SharpnessMetricsTests.swift and
RawCullTests/SharpnessScoringTests.swift.
3. Normalization And Vision
PhotoAnalysisKit normalizes the decoded CGImage to sRGB RGBA before analysis.
Vision produces attention-based saliency candidates. Classification, when
enabled, is additional metadata; it is not itself a score.
Candidates are retained when normalized area is greater than 0.03 or Vision
confidence is at least 0.9. For each retained candidate, the engine computes a
broad region detail score when it has at least 64 finite samples. Candidate
selection favors AF containment/alignment when an AF point exists, then uses
Vision confidence, detail, and area as tie-break evidence. Vision rectangles use
a bottom-origin coordinate convention; the rendered bitmap is sampled with the Y
coordinate inverted so the intended pixels are measured.
Checked against FocusMaskEngine+Scoring.swift. The package facade and
RawCull integration tests check that analysis is available; they do not directly
test candidate filtering, selection tie-breaks, or asymmetric Y-coordinate fixtures.
4. Edge-Energy Image
The package applies:
isoFactor:
ISO < 800 -> 1.0
800 <= ISO < 3200 -> 1.0 + (ISO - 800) / 2400 * 0.6
ISO >= 3200 -> min(1.6 + (ISO - 3200) / 6400 * 0.6, 2.2)
resolutionFactor =
clamp(sqrt(max(longestSide, 512) / 512), 1, 3)
effective blur radius =
min(preBlurRadius * isoFactor * resolutionFactor * apertureBlurDamp, 100)
After Gaussian pre-blur, the Metal focusLaplacian kernel computes:
laplace = 8 * center - sum(the 8 neighboring pixels)
energy = dot(abs(laplace.rgb), [0.299, 0.587, 0.114])
A color matrix multiplies that energy by 7.62. The scoring image is cropped back
to the source extent because Gaussian blur expands its extent.
Balanced and High Precision additionally build a fine pass with:
fine pre-blur = max(0.35, primary pre-blur * 0.58)
combined = primary * (1 - w) + fine * w
w = clamped to 0...0.65
Protected by: PhotoAnalysisKitTests/SharpnessMetricsTests.swift ISO and
numeric helper tests, and PhotoAnalysisKitTests/PhotoAnalyzerTests.swift
end-to-end tests.
5. Region Samples And Robust Tail Score
The engine creates these sample sets:
- full frame, excluding
borderInsetFraction on every edge; - each Vision candidate;
- a broad AF square with half-size
afRegionRadius; - AF-center and AF-neighborhood squares for evidence;
- best local patches within the AF neighborhood and winning saliency region.
Broad AF, saliency, and AF-neighborhood regions need at least 64 finite samples;
the tighter AF center needs 16. Full-frame scoring accepts any nonempty finite
sample set. Local patches use a separate patch sampling/ranking path.
For any nonempty sample array:
p20 = sample at floor((n - 1) * 0.20)
p90 = sample at floor((n - 1) * 0.90)
p97 = sample at floor((n - 1) * 0.97)
band mean = mean(max(0, value - p20))
for original samples where p90 <= value <= p97
density = min(1, (bandCount / n) / 0.06)
robust tail score = band mean * density
If p97 is not greater than p90, or the band is unexpectedly empty, the fallback
is max(0, p90 - p20). The density factor is a band-occupancy multiplier,
not an independent measure
of spatial edge density. With distinct continuous samples, the p90…p97 band
contains about 7% of samples, so the multiplier usually saturates at 1. Sparse
outliers can still score low because the top 3% is excluded and the band mean
is small; dense residual noise is not ruled out by this factor. Micro-contrast
is the standard deviation of finite Laplacian samples.
Protected by: PhotoAnalysisKitTests/SharpnessMetricsTests.swift and the
forwarding checks in RawCullTests/SharpnessScoringTests.swift.
6. Subject Construction
The broad subject score is:
AF and saliency present -> 0.6 * AF + 0.4 * saliency
only one present -> that score
neither present -> nil
The local detail score uses the same 0.6/0.4 blend for the best AF-local and
salient-interior patches. The effective subject is conservative:
broad and local present -> 0.75 * broad + 0.25 * local
only one present -> that score
This prevents one tiny high-energy patch from replacing the broad subject
measurement while still rewarding localized detail.
Checked against conservativeSubjectScore and computeSharpnessAnalysis in
FocusMaskEngine+Scoring.swift. The pinned tests do not directly assert this
blend or its missing-evidence branches.
How A Local Patch Is Chosen
“Best” means first selected by composite patch ranking, not necessarily the
patch with the largest robust-tail score. The scoring path reuses
patchRankings and selectEvidencePatches from
FocusMaskEngine+MaskGeneration.swift, applied to the scoring energy image.
Patch dimensions are 34% of the search region, bounded to 3.5%…14% of the
full image dimension, with 50% overlap. Ranking combines robust-tail score,
micro-contrast, adaptive-threshold coverage, AF proximity, an interior bonus,
silhouette penalties, and ring/compact/linear shape heuristics. AF-anchored
patches also receive a penalty for being below the AF point. A nearest-to-AF
patch is preferred when the strongest composite score is less than
1.15 times its score. The first selected patch contributes its robust-tail value
to the scalar blend.
A salient-interior patch is favored by an interior bonus; it is not required to
exclude the region border. Landscape sets the broad AF radius to zero, but it
retains the AF-neighborhood radius. An AF-local patch can therefore still
contribute to Landscape scoring. Changes to shared patch-ranking helpers can
change scalar scores even when made in the mask-generation source file.
7. Full/Subject Blend And Adjustments
When both full and subject scores exist:
weight = explicit preset override
?? aperture hint override
?? config.salientWeight
base = full * (1 - weight) + subject * weight
Two adjustments occur before the blur gate:
Silhouette penalty. The package compares average energy in the outer rim
of the effective subject region with its interior. Rim thickness is
max(1, floor(0.12 * min(regionWidth, regionHeight))) pixels, and the
comparison is borderMean / max(borderMean + interiorMean, 1e-6), rather
than the fraction of total energy located in the rim. If the derived border
fraction exceeds 0.62:
over = min(1, (borderFraction - 0.62) / 0.38)
base *= 1 - silhouettePenaltyStrength * over
Subject-size bonus. Only when no AF region exists and Vision supplied the
subject:
base *= 1 + saliencyArea * subjectSizeFactor
Fallbacks are intentionally asymmetric:
full only -> full * (1 - weight)^3
subject only -> subject
neither -> no analysis
Checked against computeSharpnessAnalysis. Existing tests cover preset values
and the analysis facade; the pinned suite has no direct regression assertions
for the cubic fallback, silhouette multiplier, or subject-size bonus.
8. Aperture-Aware Blur Gate
The effective AF analysis is preferred over saliency for subject micro-contrast.
With at least 64 samples:
t = clamp((sigma - blurGateLow) / (blurGateHigh - blurGateLow), 0, 1)
attenuation = 0.20 + 0.80 * t
final score = base * attenuation
Without a valid subject sample set, attenuation is 1.0.
Focus-failure classification is diagnostic. Here subject means the broad
saliency score, falling back to broad AF; it is not the conservative blended
subject score described above. Missing scores count as zero in the motion-blur
branch, so these labels are heuristic indicators rather than a reliable diagnosis
of the physical cause of blur:
- motion blur when global, subject, and AF-or-subject are all below 0.08 and
sigma is below 0.012;
- missed focus when global is at least 0.12 and subject/global is below
0.55;
- otherwise none.
Protected by: PhotoAnalysisKitTests/SharpnessMetricsTests.swift focus-failure
and aperture tests.
9. Score Range, UI Normalization, And Persistence
The package score is nonnegative for valid finite inputs, but the implementation
does not clamp the final score to 1.0. The subject-size multiplier can also
raise the blended value. Treat the output as a stable relative metric, not a
percentage or probability.
RawCull stores the raw value. It computes a badge denominator from the lone
score, small-set maximum, or large-set 90th-percentile element and clamps the UI
ratio to 0…1. That presentation normalization does not change the stored score
or burst-ranking input.
SharpnessAnalysisDescriptor schema 1 / algorithm 4 identifies package
behavior. It includes scoring-affecting config and policy versions, excludes
mask-only settings and per-image values, and fixes the scoring gain at 7.62.
RawCull adds source and effective pixel size, then validates source file size
and modification date on reload. Legacy descriptors decode but are stale.
Protected by:
PhotoAnalysisKitTests/SharpnessAnalysisDescriptorTests.swift;RawCullTests/SharpnessScoringTests.swift;RawCullTests/CullingModelTests.swift.
Assessment And Validation Priorities
Source review on 2026-09-28 confirms the principal formulas match the pinned
implementation. The metric is a practical relative-detail heuristic, but the
existing numeric and facade tests do not establish that its ranking is optimal
for real photographs.
Before tuning coefficients, validate these specific policies:
- Missing subject evidence: the full-only multiplier is 0.003375 for an
explicit wildlife weight of 0.85, 0.015625 for the generic default weight of
0.75, and 0.274625 for Landscape’s 0.35. Detection failure can dominate detail.
Compare a confidence-aware fallback and expose unavailable subject evidence
separately from measured softness.
- Texture and noise: band occupancy generally saturates. Evaluate sharp
low-contrast subjects, sparse real detail, blurred textured backgrounds, and
high-ISO residual noise before adding an independently measured noise or
edge-support term.
- Local subject intent: test Landscape with and without AF metadata and
test small wildlife subjects against sharper neighboring/background detail.
Decide explicitly whether AF-local scoring belongs in Landscape.
- Blur gate and scale: test ISO, aperture, thumbnail size, preview processing,
and RAW decode independently. The fixed sigma thresholds and pre-blur are
heuristics; larger thumbnails and a higher quality setting do not by themselves
prove better rankings.
- Presentation: a lone score above the denominator floor normalizes to 100%, and an all-soft
catalog can still produce Sharp labels. These are relative labels, not absolute
focus-quality judgments.
Use expert-ranked pairs from representative bursts and report pairwise ordering,
top-choice agreement, and severe subject/background mistakes by condition.
Keep evaluation images separate from coefficient tuning. Add deterministic
regressions for the confirmed failure cases, then compare proposed changes with
the current algorithm before replacing it.
Change Checklist
When changing scoring:
- change package code and package tests first, including shared patch-ranking
helpers used by scalar scoring;
- decide whether the scalar algorithm or ISO/aperture policy version must
increase;
- verify descriptor tests include every scalar-affecting setting and exclude
mask-only presentation;
- verify RawCull preset, quality, source, and size mapping;
- test batch cancellation and latest-generation publication;
- inspect raw score distributions separately from normalized badge labels;
- update this reference and the overview together.
7 - Detailed Focus Mask Computation
Detailed Focus Mask Computation
This page follows visible focus-mask rendering at PhotoAnalysisKit 1.3.1,
revision 2a1466e04d821fa2628d6985296643e0d0c7e465.
The focus mask is not the camera’s AF-point marker. The marker reports where the
camera attempted focus. The mask is a package-generated bitmap of selected
edge-detail evidence. The AF point can guide selection, but it is never painted
as the mask by itself.
RawCull owns the SwiftUI trigger, image and metadata supplied to the package,
task lifetime, and overlay presentation. PhotoAnalysisKit owns normalization,
saliency, evidence selection, Laplacian generation, patch ranking, thresholding,
morphology, colorization, clipping, and mask diagnostics.
For shared ownership, persistence, calibration, and scalar score behavior, see
Focus Mask And Sharpness.
Data Shapes And The Overlay Boundary
flowchart LR
VIEW["SwiftUI view<br/>NSImage or CGImage"] --> FM["@MainActor FocusMaskModel"]
FM --> INPUT["PhotoAnalysisInput<br/>CGImage<br/>ISO, aperture, AF CGPoint?"]
INPUT --> PA["PhotoAnalyzer"]
PA --> NORM["normalized sRGB CGImage"]
NORM --> CI["CIImage RGBAf analysis"]
CI --> VALUES["Float edge samples<br/>saliency rectangles<br/>FocusPatchRanking[]"]
VALUES --> EVIDENCE["FocusEvidence<br/>regions, confidence, threshold, coverage"]
VALUES --> MASKCI["thresholded/colorized CIImage"]
MASKCI --> OUT["CGImage focus mask"]
OUT --> ADAPT["NSImage when required"]
ADAPT --> OVERLAY["SwiftUI Image overlay<br/>presentation boundary"]
EVIDENCE --> BREAKDOWN["SharpnessBreakdown diagnostics"]CGImage is the public decoded-image boundary. Package internals create
CIImage values and render RGBAf edge-energy samples with a reusable
CIContext. The returned CGImage contains only the overlay bitmap; RawCull
positions it over the displayed image. The package does not retain a SwiftUI
view, app model, URL, or persistence object.
Call Paths
| RawCull context | Model call | Package facade | Evidence behavior |
|---|
| Main thumbnail/detail | generateFocusMask | PhotoAnalyzer.focusMask | Can reuse existing FocusEvidence; classification is skipped and a saved winning saliency rectangle avoids a new saliency pass |
| Zoom overlay | generateFocusMaskWithBreakdown | PhotoAnalyzer.analyzeWithFocusMask | Recomputes saliency, scalar breakdown, evidence, and mask together |
| Comparison grid | generateFocusMaskWithBreakdown | PhotoAnalyzer.analyzeWithFocusMask | Same aligned diagnostic path per comparison image |
The RawCull entry files are:
Views/ThumbnailComponents/MainThumbnailImageView.swift;Views/ZoomViews/ZoomOverlayView.swift;Views/ComparisonGridView/ComparisonGridImageCoordinator.swift;Model/ViewModels/FocusandSharpness/FocusMaskModel.swift.
Views snapshot the effective config, inject ISO, aperture, and normalized AF
point, cancel a previous mask task when image or config identity changes, and
check cancellation before publishing the result.
Which Values Affect What
The categories below are important when tuning the mask.
Scalar Score Values
These change package scalar analysis and therefore belong in
SharpnessAnalysisDescriptor:
- pre-blur radius;
- border inset;
- salient weight and explicit override;
- subject-size factor;
- silhouette-penalty strength;
- broad/center/neighborhood AF radii used by scoring evidence;
- fine-detail blend;
- classification policy;
- the stable scoring gain and algorithm/policy versions.
ISO and aperture also change scalar analysis, but are per-image inputs rather
than descriptor configuration. Photo type and quality change scalar output
indirectly by applying the values above.
Mask Presentation Values
These change only the rendered overlay or its visibility:
| Value | Current default | Effect |
|---|
threshold | 0.46 | Fallback/floor reference for adaptive visual threshold |
dilationRadius | 0.0 | Optionally connects thresholded edge pixels |
erosionRadius | 0.0 | Optionally removes isolated/small responses |
featherRadius | 0.5 | Softens the clipped mask alpha |
showRawLaplacian | false | Debug early return before threshold, patches, color, and morphology |
guaranteeVisibleFocusEvidence | false | Compatibility property; no longer lowers the render threshold |
minimumEvidenceCoverage | 0.001 | Compatibility property retained with the public configuration |
isolateMaskToSubject | true | Chooses subject/AF search regions rather than the whole frame |
The final mask uses a single warm red/orange color matrix: red 1.0, green 0.22,
blue 0.02, alpha 0.92 before SwiftUI compositing.
Shared Evidence Values
Some values affect both analysis evidence and rendering even though the final
mask threshold itself never becomes a scalar multiplier:
- pre-blur radius, energy gain, ISO, aperture blur damp, and image resolution
determine the shared primary Laplacian;
- border inset excludes scalar border samples and blackens the corresponding
mask border;
- saliency candidates and AF point determine scoring regions and potential mask
search regions;
- AF radii determine scalar evidence regions and which mask region can be
selected;
- existing
FocusEvidence can make mask selection follow a previous score.
Calibration Output
Calibration returns
FocusCalibrationResult(threshold, sampleCount, p50, p90, p95, p99). RawCull
applies only threshold to the active config. The percentiles and sample
count are diagnostics. Calibration changes visual threshold policy; it does not
rescale stored scalar scores.
Stage 1: Normalize And Resolve Per-Image Configuration
PhotoAnalyzer.focusMask and analyzeWithFocusMask normalize the input
CGImage to sRGB. They copy the caller’s SharpnessConfiguration, then
replace:
config.iso = input.iso
config.apertureHint = derived from input.aperture
Aperture at or below f/5.6 is wide, at or above f/8 is landscape, the interval
between is mid, and missing metadata defaults to mid.
The package engine runs synchronous Vision/Core Image work in an explicitly
@concurrent child task. A cancellation handler cancels that worker. This is
not a detached task in the pinned revision.
Stage 2: Obtain Saliency And Score Evidence
The simple facade follows this rule:
saved winningSaliencyRect exists -> reuse it
else isolateMaskToSubject -> run saliency, without classification
else -> no salient region
The diagnostic facade always detects saliency, optionally classifies, computes a
SharpnessBreakdown, then gives the resulting evidence to mask rendering. That
keeps the displayed patch, score diagnostics, and saliency decision from the
same analysis pass.
The evidence can request one of:
- AF center;
- AF neighborhood;
- broad AF point;
- saliency;
- mixed AF and saliency;
- global;
- none.
If the requested evidence is unavailable, rendering falls back in this order:
broad AF, saliency, then global.
Stage 3: Scale And Build The Primary Laplacian
The input CIImage is transformed by the requested mask scale. RawCull’s
FocusMaskAnalysisResolutionPolicy.prepare returns every decoded preview pixel,
so view code must choose an appropriate decode before this call. The package
builds a dedicated native-pixel focus-mask detail image with clamped edges and
the following pre-blur:
focus-mask pre-blur = max(0.35, primary preBlurRadius * 0.52)
Native-mask mode disables resolution scaling and preserves the input extent.
The scalar scoring path continues to use its resolution-aware primary and
optional fine-detail passes.
The underlying scalar edge pipeline is:
ISO factor:
below 800 1.0
800...3199 linear 1.0 -> 1.6
3200 and above linear tail capped at 2.2
resolution factor =
clamp(sqrt(max(longestSide, 512) / 512), 1, 3)
blur radius =
min(preBlurRadius * ISO factor * resolution factor * aperture damp, 100)
Laplacian =
abs(8 * center - 8 neighbors)
collapsed with [0.299, 0.587, 0.114]
multiplied by energyMultiplier
Landscape aperture damp is 0.8; wide and mid are 1.0. The package then blackens
the configured border inset so expanded Gaussian edges are not visual evidence.
Raw Laplacian Debug Mode
When showRawLaplacian is true, the package crops the boosted Laplacian to the
scaled image extent and returns it immediately. It records the saliency/AF
region source, but does not:
- build or select patches;
- calculate an adaptive visual threshold;
- threshold or colorize edges;
- erode, dilate, thin, clip, or feather the mask.
Use this mode to inspect the edge-energy input, not the final overlay.
Stage 4: Resolve Search Regions
With subject isolation enabled, the package converts normalized Vision and AF
coordinates into pixel rectangles. AF Y is inverted when moving between the
view/AF convention and Core Image pixel space.
It also creates two tighter AF rectangles from defaults:
AF center half-radius 0.025 of image dimensions
AF neighborhood half-radius 0.075
broad AF half-radius config.afRegionRadius
Search regions are selected from the requested evidence. Mixed evidence searches
AF and saliency separately. Global/none searches the whole image. With subject
isolation disabled, the whole image becomes the saliency-shaped selection and
the search is effectively global.
Every focus-mask region uses the same native-pixel detail source described in
Stage 3. The mask is not center-weighted around the AF point. This 0.52 factor
is mask selection and rendering policy; scalar quality’s second pass uses a
different 0.58 factor and blend.
Stage 5: Generate Candidate Patches
For each bounded search region:
patch width =
min(max(region width * 0.34, image width * 0.035), image width * 0.14)
patch height =
min(max(region height * 0.34, image height * 0.035), image height * 0.14)
grid step = patch dimension * 0.50
If the AF-centered patch overlaps at least 75% of its intended size, it is
added. The renderer also appends a separate 6%-of-image AF patch when an AF
point exists.
Each patch records:
- robust p90…p97 tail detail relative to p20;
- micro-contrast standard deviation;
- coverage above an adaptive patch threshold;
- normalized distance to AF and whether it contains AF;
- interior versus silhouette fraction;
- ring, compact-detail, and linear-edge shape evidence;
- a below-AF penalty and eye/head heuristic adjustment.
Stage 6: Rank Patches
The current composite is:
robust tail
+ micro contrast * 0.35
+ coverage * 0.08
+ AF proximity * 0.12
+ interior bonus (0.03 when not touching search border)
- silhouette * (0.18 for AF-anchored regions, otherwise 0.45)
+ eye/head adjustment
eye/head adjustment =
ring detail * 0.10
+ compact detail * 0.08
- linear edge * 0.10
- below-AF penalty
below-AF penalty =
clamp((distance below AF - 0.025) / 0.15, 0, 1) * 0.18
AF proximity falls linearly to zero at normalized distance 0.20. Composite
scores are floored at zero.
For AF-anchored evidence, the nearest patch is promoted ahead of the strongest
only when it is not already strongest and the strongest is less than 1.15 times
the nearest. Selection then walks descending composite score, rejects patches
with overlap ratio 0.55 or greater, and keeps at most three.
These rankings are visual-evidence selection and diagnostics. Scalar scoring
does use the best AF-local and salient-interior patch’s robust tail score as
a conservative 25% refinement of broad subject evidence, but it does not use the
rendered mask threshold, morphology, color, or rendered coverage.
Stage 7: Choose The Visual Threshold
Samples are collected from the full selected search regions. Ranked patches
summarize and order local evidence, but no longer truncate the visible focus
map. The adaptive threshold is:
| Evidence | Percentile | Floor from config threshold | May exceed fallback? |
|---|
| AF center/neighborhood/point | 0.82 | 0.32 times fallback | No |
| Saliency, mixed, or global | 0.90 | 0.55 times fallback | Yes |
The floor is at least 0.01 and the final threshold is at most 0.95.
The threshold is not relaxed for visibility. relaxedForVisibility is always
false in the current renderer, and weak images may return an empty mask.
Stage 8: Render And Clip
The package performs this sequence:
- copy the Laplacian red channel into grayscale;
- apply
CIColorThreshold; - apply morphology minimum with
erosionRadius, when positive; - apply morphology maximum with
dilationRadius, when positive; - colorize to warm red/orange with alpha 0.92;
- clip the result to the union of the selected search regions;
- Gaussian-feather with
featherRadius, when positive; - crop to the scaled image extent and create the output
CGImage.
After morphology and feathering, GPU area reductions compute visible-alpha
coverage within the selected regions. This final rendered coverage—not a
pre-render sample estimate—is stored in the evidence.
Stage 9: Return Diagnostics
The mask result updates FocusEvidence with:
- visualized region and overlay style;
- all sorted patch rankings;
- effective threshold, coverage, and relaxation flag;
- visualized centroid distance from AF;
- spatial-alignment score and best/second-best dominance;
- silhouette indication;
- high, medium, or low evidence confidence plus a reason.
Confidence rules are intentionally readable:
- no viable patch -> low;
- rendered coverage below 0.001, robust-tail below 0.01, micro-contrast below
0.005, or patch coverage below 0.001 -> low;
- AF anchored within normalized distance 0.05 with strong detail -> high;
- AF aligned with measurable but weaker detail -> medium;
- AF farther than 0.05 -> low;
- global with composite at least 0.10 -> medium, otherwise low;
- non-AF subject with strong detail, silhouette below 0.20, and dominance at
least 1.08 -> high;
- other usable subject evidence -> medium.
The diagnostic facade also records FocusMaskRegionSource and the visual
threshold in SharpnessBreakdown. RawCull adapts that package breakdown only to
add the selected scoring source.
SwiftUI Publication And Cancellation
RawCull views own their mask tasks:
- a new image, config, source, or explicit regeneration cancels the old task;
- view disappearance/toggle-off cancels and clears mask state;
- package workers check cancellation before and after Vision, Laplacian, patch,
and render stages;
- views check cancellation before assigning the returned mask and breakdown.
The package result is a value. It cannot publish into an old view on its own;
the owning SwiftUI task is the final stale-result boundary.
Debugging A Surprising Mask
- Confirm the input image, mask scale, ISO, aperture, and normalized AF point.
- Inspect
winningRegion, winning saliency rectangle, and region source. - Compare AF-center, AF-neighborhood, broad AF, saliency, and global scores.
- Inspect sorted patch composite components, especially AF distance,
silhouette, linear-edge, and below-AF penalties.
- Check effective visual threshold and final rendered coverage.
- Enable raw Laplacian mode to separate edge-energy input from threshold and
morphology.
- Remember that a strong scalar score and a sparse mask are compatible: the
mask is permitted to be empty when the evidence gates are not met.
Protecting Tests
| Stage | Tests |
|---|
| RawCull facade, Metal resource, returned breakdown, scoring-source adaptation | RawCullTests/PhotoAnalysisKitIntegrationTests.swift |
| Public analyze/mask/calibration behavior and cancellation | PhotoAnalysisKitTests/PhotoAnalyzerTests.swift |
| Native-pixel mask detail, full-region rendering, empty weak masks, and final coverage | PhotoAnalysisKitTests/FocusMaskAccuracyTests.swift |
| Robust tail, micro-contrast, ISO curve, aperture gates, failure classification, presets | PhotoAnalysisKitTests/SharpnessMetricsTests.swift |
| Descriptor excludes mask-only presentation and includes scalar policy | PhotoAnalysisKitTests/SharpnessAnalysisDescriptorTests.swift |
| Bounded input loading and analysis cancellation | PhotoAnalysisKitTests/PhotoAnalysisBatchTests.swift |
Change Checklist
When changing the mask:
- classify the value as scalar-affecting, mask-only, shared evidence, or
calibration output;
- update the descriptor only for scalar-affecting behavior;
- verify simple rendering can reuse score evidence without repeating Vision;
- verify AF and Vision coordinate conversion;
- test cancellation at the package worker and SwiftUI publication boundaries;
- inspect raw Laplacian, patch diagnostics, threshold coverage, and final
overlay as separate stages;
- update this page and the overview together.
8 - Burst Groups
Burst Groups
Burst analysis turns a flat catalog into groups of adjacent, visually similar
frames. It ranks each multi-frame group, presents review queues, and supports
culling decisions such as keeping the best frame, keeping the top two, deferring
a group, or setting a manual pick.
The implementation separates application commands in RawCullViewModel, worker
orchestration in BurstAnalysisCoordinator, backend-selectable similarity
indexing behind RawCullSimilarityFeature, pure grouping and ranking engines,
two levels of artifact persistence, and the review UI. Vision feature prints are
the safe default. A validated CLIP model can become the active burst-similarity
backend when the user enables it; the rest of the burst pipeline works with
typed SimilarityArtifact values rather than assuming one representation.
Source Map
| Area | Main files |
|---|
| Application commands and result publication | RawCull/Model/ViewModels/RawCullViewModel+BurstGrouping.swift |
| Worker orchestration and cache compatibility | RawCull/Intelligence/BurstAnalysis/BurstAnalysisCoordinator.swift and coordinator extensions |
| Similarity feature and shared state | RawCull/Intelligence/Similarity/RawCullSimilarityFeature.swift, SimilarityScoringModel.swift |
| Backend composition and adapters | RawCull/Intelligence/Composition/RawCullAIIntegration.swift, RawCull/Intelligence/Similarity/RawCullVisionSimilarityService.swift |
| Per-file durable artifacts | RawCull/Intelligence/Persistence/PerFileAnalysisArtifactStore.swift |
| Pure grouping and ranking | RawCullCore Sources/RawCullCore/BurstGroupingEngine.swift, BurstRankingEngine.swift |
| Shared models | RawCullCore Sources/RawCullCore/BurstAnalysisModels.swift; app BurstAnalysisModels.swift, BurstReviewQueueModels.swift |
| Cache and repository boundary | RawCull/Intelligence/Persistence/BurstAnalysisCache.swift, RawCull/Intelligence/BurstAnalysis/BurstAnalysisCacheRepository.swift |
| Ratings and manual overrides | CullingModel.swift, SavedFiles.swift |
| Burst home and review list | BurstGroupsHomeView.swift, SimilarityGridSelectionView.swift, CullingGridView.swift |
| Single-burst workspace and comparison | BurstCullingWorkspaceView.swift, ComparisonGridView.swift |
| Batch badge selection and rating | CullingGridSelectionCoordinator.swift, CullingGridView.swift |
| Deep Review subject outlines | DeepAIReviewMaskOutlineRenderer.swift, MainThumbnailImageView.swift, ZoomOverlayView.swift, BurstCullingWorkspaceView.swift |
| Tests | RawCullCore/Tests/RawCullCoreTests/BurstGroupingEngineTests.swift, BurstRankingEngineTests.swift, app burst/culling tests |
End-to-End Flow
flowchart TD
A["Analyze Bursts"] --> B["RawCullViewModel builds immutable request"]
B --> C["BurstAnalysisCoordinator owns generation and task"]
C --> H0["Hydrate valid per-file similarity artifacts"]
H0 --> D{"Valid BurstAnalysisCache and matching artifact digest?"}
D -->|"yes"| E["Remap cached IDs and apply snapshot"]
D -->|"no"| F{"Sharpness scores missing?"}
F -->|"yes"| G["Calibrate and score target files"]
F -->|"no"| H["Reuse scores"]
G --> I{"Similarity artifacts missing?"}
H --> I
I -->|"yes"| J["Index with active Vision or CLIP service"]
I -->|"no"| K["Reuse descriptor-valid artifacts"]
J --> L["Commit per-file artifacts"]
K --> M["Group adjacent frames"]
L --> M
M --> N["Rank multi-frame groups"]
N --> O["Apply manual winner overrides"]
O --> P["Save derived snapshot and review states"]
P --> Q["Show dashboard, queues, and workspace"]The coordinator owns the worker generation, task, progress, cache preparation,
missing sharpness/similarity work, grouping, ranking, and the primary cache
save. It receives callbacks for the application-owned catalog validity check and
final result publication. Every awaited phase is therefore protected by both the
coordinator generation and selected catalog; a cancelled or superseded run
cannot publish late results into a newer catalog.
The target is normally every catalog file sorted by localized filename, which
acts as shot order. If files are selected, visible selected files are followed
by hidden selected files. If no selection exists and a star filter is active,
only files with that rating are analyzed.
Similarity Artifacts And Backend Selection
SimilarityScoringModel depends on the RawCullSimilarityServicing protocol.
The active service supplies a backend descriptor, the descriptors it can
produce, an indexing operation, and a distance operation. This keeps grouping
independent of whether the payload is a Vision feature print or a CLIP image
embedding.
RawCullAIIntegration is the composition root:
RawCullVisionSimilarityService is always available and is the
startup/default service.RawCullCLIPSimilarityService is selected only when CLIP is enabled and the
chosen model bundle has validated and produced a provider.- The CLIP service can recover through its configured Vision provider.
Diagnostics record partial CLIP generation and whole-batch Vision fallback
instead of silently changing artifact meaning.
| Setting | Current value |
|---|
| Input thumbnail maximum | 512 px |
| Similarity pipeline version | 3 |
| Artifact schema | SimilarityArtifactDescriptor.currentSchemaVersion |
| Default backend | Vision feature print |
| Optional backends | Validated OpenAI or DataComp CLIP provider |
Each SimilarityArtifact contains a descriptor and encoded payload. The
descriptor records backend identity, model fingerprint, representation,
preprocessing, normalization, configuration, and schema version.
RawCullSimilarityArtifactValidation compares that descriptor and the source
fingerprint before an artifact is admitted.
PerFileAnalysisArtifactStore persists individually valid artifacts
independently of the catalog-wide burst snapshot. On a later run,
hydrateArtifacts(_:) loads only artifacts allowed by the current service and
pipeline signature. This makes a partial index reusable and lets invalid entries
be removed without discarding every other file.
The same model can rank a catalog by distance from an anchor image. Burst
grouping calculates distances only between adjacent files. Those distances are
cached in memory under the current artifact/backend signature and reused when
regrouping remains compatible.
Grouping Rules
BurstGroupingEngine.group(...) makes one sequential pass. It starts a new
group when any boundary rule fires.
| Boundary reason | Trigger |
|---|
| Visual distance changed | Adjacent active-backend distance is at or above visualDistanceThreshold |
| Similarity evidence missing | No adjacent distance is available |
| Capture gap | Absolute capture-date gap exceeds maxTimeGapSeconds; modification-date fallback uses maxFallbackTimeGapSeconds |
| Camera changed | Normalized camera value changed and requireSameCamera is enabled |
| Focal length changed | Parsed focal-length delta exceeds maxFocalLengthDeltaMM |
| Exposure changed | Aperture changes by more than 0.2, ISO changes, or shutter-speed text changes |
Lens changes are recorded as evidence but do not independently split a group.
They do make group metadata unstable during ranking.
Default configuration:
visualDistanceThreshold = 0.25
maxTimeGapSeconds = 2.0
maxFallbackTimeGapSeconds = 10.0
requireSameCamera = true
requireSimilarFocalLength = true
maxFocalLengthDeltaMM = 3.0
algorithmVersion = 4
The burst sensitivity control changes only the visual threshold.
reGroupBursts() cancels older grouping work, reuses similarity artifacts and
adjacent-distance data, rebuilds rankings, and saves a new cache. The current
home/category presentation is preserved instead of being forced into the grouped
grid.
BurstRankingEngine computes:
overall =
rankingSharpness * 0.62
+ focusPoint * 0.12
+ saliency * 0.10
+ metadata * 0.16
Sharpness is normalized by SharpnessScoringModel.maxScore. When at least two
group members have scores and their normalized spread is at least 0.03, global
and burst-relative sharpness are blended:
rankingSharpness = normalizedSharpness * 0.65
+ burstRelativeSharpness * 0.35
The other components are heuristic evidence:
- Focus is 0.70 when camera AF data exists and 0.45 otherwise.
- Saliency is 0.75 when the subject label matches the group’s dominant label,
0.25 on a mismatch, and 0.45–0.60 when evidence is incomplete.
- Metadata starts from group stability, gains 0.15 for tight similarity, loses
0.10 at ISO 6400 or above, and gains 0.05 at f/5.6 or wider.
Ties in overall score retain original shot order.
Confidence And One-Click Safety
| Confidence | Conditions |
|---|
| High | Scores exist, group has at least 3 files, best leads second by at least 0.12, best normalized sharpness is at least 0.65, metadata is stable, and all internal visual distances are below 0.22 |
| Medium | Best leads by at least 0.05 and metadata is stable |
| Low | Scores are absent, candidates are close, or evidence is unstable |
isSafeForOneClickCulling is true only for high-confidence results.
keepBestInGroup and keepTopTwoInGroup also require the current sharpness
score table to be non-empty; otherwise they return without changing ratings.
Home Dashboard And Review Queues
After analysis, BurstGroupsHomeView shows catalog coverage, group counts, the
active similarity threshold, up to three suggested picks, and these queue
categories:
| Category | Selection rule |
|---|
| All | Every computed group, including singleton groups |
| Single Images | Groups containing exactly one file |
| Needs Review | Multi-frame groups with explicit review-needed state or unsafe/uncertain ranking evidence |
| Deferred | Multi-frame groups explicitly deferred |
| Marked Reviewed | Multi-frame groups explicitly marked .reviewed |
| Reviewed | Effective reviewed results, decisions already applied, and manual-winner groups |
The grouped culling grid can collapse a burst to its top three ranked frames. A
group header opens the dedicated workspace and toggles Reviewed or Deferred
state.
The grid can also derive batch selectors from visible burst-rank, saliency, and
sharpness badges. A normal badge action replaces the selection with every
visible match, Command toggles the matching set, and Shift extends/replaces
according to the coordinator’s modifier policy. Batch rating applies one rating
to the resulting selected files. The coordinator is a pure value transformation
so these semantics are testable without SwiftUI.
BurstCullingWorkspaceView displays one large selected frame plus a bounded
three-frame image window around the current selection and a filmstrip of ranked
candidates. It reuses ComparisonImagePaneView,
ComparisonViewportInteractionState, ImageSourceSelectionState, and
ZoomMetadataPanel rather than creating a second image-inspection
implementation. The cache key combines file identity and selected image source.
P/N and the arrow keys move between frames, G advances to the next
eligible multi-frame group, and E toggles the metadata panel. The workspace
also exposes zoom, thumbnail/embedded-JPEG source selection, focus evidence,
rating, pick/reject, reviewed state, and the detailed comparison grid. Escape
returns to the active burst list.
When Deep Review has produced a stored mask for the selected file, the workspace
can render its orange subject outline. S toggles the outline. The lookup uses
completed mask candidates indexed by file ID, and asynchronous outline results
are committed only while the selected file and mask identity remain current.
Review States
RawCullCore.BurstReviewState currently defines:
.none.needsReview.reviewed.deferred.algorithmReviewed.manualWinnerOverride.decisionApplied
The app and package enum are now aligned. algorithmReviewed remains for cache
compatibility.
BurstReviewQueuePolicy.effectiveState(for:) preserves explicit modern states.
For .none or legacy .algorithmReviewed, it derives Needs Review when
confidence is not high, cautions exist, no recommendation exists, or one-click
culling is unsafe; otherwise it treats the group as reviewed.
Toggling Reviewed or Deferred a second time resets the result to .none and
removes that group from the explicit state dictionary. Review states are
persisted using a catalog-and-membership BurstGroupSignature, so they can be
restored after a threshold change even if numeric group IDs change.
Manual Winner Overrides
A manual winner is stored in savedfiles.json as BurstWinnerOverride:
| Field | Meaning |
|---|
winnerFileName | User-selected winner |
memberFileNames | Group membership snapshot |
Overrides use filenames because saved-file persistence is filename-based.
CullingModel canonicalizes member names so lookup is order-independent and
prunes overrides whose files no longer exist. Applying an override promotes the
selected file, recalculates second place, and sets .manualWinnerOverride.
User Actions
| Action | Method | Effect |
|---|
| Keep best | keepBestInGroup | On a safe result, rate the winner 3 stars and reject the rest |
| Keep top two | keepTopTwoInGroup | On a safe result, rate first place 3 stars, second place 2 stars, and reject the rest |
| Set manual pick | setManualBurstWinner | Persist the winner override and rate the selected frame 3 stars |
| Open group | compareBurstGroup | Open the workspace with up to four ranked comparison IDs |
| Next group | advanceToNextBurstGroup | Open the next eligible multi-frame group in the active queue |
| Toggle reviewed/deferred | toggleBurstGroupReviewed, toggleBurstGroupDeferred | Persist or clear the explicit review state |
| Undo last burst action | undoLastBurstAction | Restore the previous ratings captured for the last one-click action |
| Reindex | reindexBurstAnalysis | Clear loaded analysis, delete the catalog cache, and recompute |
One-click rating actions capture a BurstUndoEntry before writing ratings. Only
the most recent burst action is retained for undo.
Cache Validity
BurstAnalysisCache stores similarity artifacts, scores, saliency, groups,
boundary evidence, ranked results, and review-state snapshots. It is a derived
catalog snapshot; PerFileAnalysisArtifactStore is the reusable per-image
artifact layer. The current burst-cache schema is 9. A snapshot is accepted only
when all of these still match:
- cache schema version,
- grouping algorithm version,
- catalog path,
- effective sharpness thumbnail size and complete sharpness signature,
- grouping configuration and active backend descriptor,
- all allowed artifact backend descriptors, artifact schema, input size, and
pipeline version,
- file count,
- every file path, size, and modification date,
- a digest of the current descriptor-and-payload artifact set.
On load, every embedded artifact is revalidated against its source and allowed
backend descriptors. Cached UUIDs are then remapped to the current scan’s UUIDs
by file path because FileItem.id values are recreated. Cache saves are guarded
by the completed analysis context, generation, catalog, artifact digest, and
similarity signature so stale asynchronous work cannot overwrite a newer result.
Schema 8 can be read only as a migration candidate; individually valid artifacts
and stable review-state signatures may be imported, but the old snapshot is not
treated as a current cache hit.
What To Check When Changing This Area
- Bump
BurstGroupingConfig.algorithmVersion when grouping semantics change. - Update the similarity pipeline version or descriptor/signature when artifact
inputs or meaning change.
- Preserve descriptor validation and the separation between per-file artifacts
and the derived burst snapshot.
- Test both the always-available Vision path and validated CLIP
selection/fallback behavior.
- Update the sharpness scoring signature when score meaning changes.
- Keep review-state decoding backward compatible and preserve signature-based
restoration across regrouping.
- Keep manual winner overrides durable across cache invalidation and UUID
remapping.
- Treat high confidence plus available sharpness scores as the one-click culling
gate.
- Keep the workspace’s bounded image window and source-aware cache identity when
changing image navigation.
9 - Future Features and Competitive Evaluation
Competitive feature review and a prioritized roadmap for RawCull after version 3.2.2, including an evaluation of additional Core AI models.
Future Features and Competitive Evaluation
This page evaluates possible RawCull features after version 3.2.2. It compares
the current implementation with representative professional culling products,
identifies the most important workflow gaps, and evaluates additional models
from Apple’s Core AI model catalog.
The comparison was researched on 1 September 2026. Competitor products and
Apple’s model catalog change frequently, so links to primary vendor sources are
included. The comparison is based on documented functions and inspection of
the RawCull source; it is not a controlled image-quality or speed benchmark.
The term leading applications in this document means representative products
with significant professional culling functionality. It is not a market-share
ranking.
Executive Recommendation
RawCull should remain a focused, local, explainable culling application. It
should not try to become a complete editor, retouching suite, cloud gallery, or
delivery platform.
The most valuable next release would combine three features:
- XMP interoperability with Lightroom Classic and Capture One;
- face and eye inspection for portraits, events, and group photographs; and
- an explainable, non-destructive first pass that organizes photographs into
Keep, Review, and Likely Reject.
A suitable product description would be:
RawCull First Pass: private, local, and explainable culling with face and eye
inspection and seamless XMP handoff.
The first additional Core AI model to evaluate should be EfficientSAM. The
RawCull and PhotoAIKit integration already contains an EfficientSAM backend,
and the model is much smaller than SAM 3. Object detection and a small
vision-language model are reasonable later experiments, but they should not be
added until a measured culling problem justifies their download, memory, and
maintenance costs.
Product Position
RawCull’s strongest differentiator is not simply that it uses AI. Several
competitors also run culling locally. Its differentiator is the combination of:
- fast embedded RAW previews;
- Sony and Nikon camera autofocus-point extraction;
- full-frame, salient-subject, and AF-region sharpness evidence;
- visible focus masks and an explanation of ranking cautions;
- local visual similarity and burst grouping;
- local CLIP semantic search;
- non-destructive rating and review decisions; and
- no requirement to upload photographs for analysis.
This is a credible position: show the photographer why one frame is stronger,
and leave the decision with the photographer.
Future automation should preserve that contract. RawCull should not silently
delete a file, conceal a low-confidence result, or present a generative
explanation as measured fact.
Version 3.0.0 And Current 3.2.2 Baseline
Version 3.0.0 is sometimes described as the non-AI version. More precisely, it
does not require separately downloaded AI models. It still uses Apple Vision
feature prints for image similarity and burst grouping, and PhotoAnalysisKit
uses Vision and Metal-based analysis. It does not provide CLIP text-to-image
search or segmentation-backed Deep Review.
The current 3.2 line adds the optional Core AI layer and later 3.2.2 work
completes the SAM 3 release path, subject outlines, and batch grid selection:
| Capability | Version 3.0.0 | Version 3.2.2 |
|---|
| Embedded-preview culling | Yes | Yes |
| Sony ARW, Nikon NEF, and Adobe DNG | ARW and NEF | Yes |
| EXIF and camera AF point | Yes | Yes |
| Sharpness calibration and focus masks | Yes | Yes |
| Vision feature-print similarity | Primary backend | Fallback backend |
| Burst grouping and candidate ranking | Yes | Yes |
| Local CLIP similarity | No | Optional |
| Natural-language semantic search | No | Optional; requires CLIP |
| SAM 3 Deep Review | No | Production-enabled, optional download |
| Managed model download workflow | No | Yes |
| Cached Deep Review subject outlines | No | Loupe, zoom, burst workspace, and review sheet |
| Badge-based batch selection/rating | No | Yes |
The normal culling workflow is deliberately shared. Features such as XMP,
improved ingest, broader formats, better review queues, and keyboard workflow
can therefore benefit both release lines. Model-dependent features should
degrade cleanly when no model is present.
Current RawCull Capabilities
Source inspection confirms the following current application behavior:
- catalog discovery for registered Sony ARW, Nikon NEF, and Adobe DNG files;
- concurrent EXIF, dimensions, camera, lens, ISO, aperture, and AF metadata;
- two-tier thumbnail caching plus full-size embedded/developed preview caches;
- AF overlays and GPU-generated focus masks;
- configurable sharpness scoring using full-frame, salient-subject, and
AF-region evidence;
- visual grouping of neighboring frames and ranked burst candidates;
- manual comparison, burst workspaces, review/defer state, and manual winners;
- local CLIP image embeddings and text-query embeddings;
- Vision fallback when CLIP is unavailable or disabled;
- optional SAM 3 subject-mask Deep Review, including multi-subject union masks
and cached orange subject outlines;
- badge-based batch selection and rating in the culling grid;
- reject, neutral keeper, and two- through five-star rating states;
- persisted ratings, analysis artifacts, burst decisions, cache signatures,
and settings;
- embedded or developed JPEG export;
- rsync-based copying of tagged or minimum-rated RAW files, with progress and
cancellation; and
- memory-pressure monitoring and live cache accounting in Settings.
The application currently stores its rating decisions in RawCull’s own JSON
data rather than writing standard XMP sidecars. This protects source files but
limits interoperability.
Competitive Products Reviewed
Aftershoot
Aftershoot offers automated, assisted, and manual culling. Its documented
workflow includes duplicate grouping, key faces, blur and closed-eye detection,
image scores, Selected/Highlight/Maybe/Blur/Closed Eyes groups, Survey Mode,
strictness controls, and preserved ratings and labels when exporting to
Lightroom Classic or Capture One.
Aftershoot also describes local/offline culling and non-destructive XMP
sidecars for RAW photographs. Its 2026 material promotes selection-style
learning and culling to a requested yield, although individual functions can
have staged availability.
Sources:
Narrative Select
Narrative Select emphasizes assisted rather than opaque automatic decisions.
Its notable functions are synchronized face close-ups, individual face focus
and eye-state assessments, scene grouping, first-pass indicators, image
sharpness ranking, a people filter, rapid RAW loading, and direct handoff to an
editing workflow.
Narrative documents support for Canon CR2/CR3/CRW, Nikon NEF/NRW, Fuji RAF,
Sony ARW, Panasonic RW2, Olympus ORF, Pentax PEF, DNG, generic RAW, JPEG, HEIC,
and HEIF. It also states that images remain local and are not uploaded for
processing unless the user separately opts in to improvement use.
Sources:
Adobe Lightroom Classic
Lightroom Classic Assisted Culling can select by subject focus, eye focus, and
eyes open. It can reject documents, receipts, probable misfires, and exposure
problems, and it can auto-stack by time and visual similarity. Results can be
turned into flags, ratings, labels, folders, or collections.
The June 2026 Faces panel presents every detected face together with eye-open
and eye-sharpness scores. Lightroom also added catalog-wide exact duplicate
detection across folders. This integration is important because it removes the
handoff step entirely for existing Lightroom customers.
Sources:
Imagen
Imagen offers two clear automated policies: retain the best photograph from
each similar group, or cull to an exact number or percentage. It can identify
duplicates and low-rated images with blur, accidental-capture, or exposure
problems. Users review the results before sending selected photographs into
Imagen’s editing workflow.
Imagen can show an edited preview before the culling decision, but culling is
performed on Imagen Cloud. This is a different product tradeoff from RawCull’s
fully local analysis.
Sources:
FilterPixel
FilterPixel separates fast technical filtering from a genre-aware DeepCull.
The documented criteria include blur, focus, expressions, composition,
lighting, background, narrative value, brand safety, emotional moments, and
peak sports action. It promotes Keep/Review/Reject groups, best-of-burst
selection, target counts, and a score plus reason for each choice.
FilterPixel’s public pages observed during this review were inconsistent about
whether DeepCull processing is local or cloud-based. RawCull should therefore
not use FilterPixel’s processing location as a competitive fact without a new
verification. Its genre-aware selection and reason presentation are still
useful product references.
Sources:
Photo Mechanic
Photo Mechanic remains a useful benchmark for ingest and metadata rather than
automatic AI selection. Its ingest can copy in the background, rename files,
apply metadata templates, create folders, track copied photographs, and open a
contact sheet while copying. It also provides tags, stars, configurable color
classes, XMP/IPTC interoperability, variables, and code replacements.
Sources:
Capability Matrix
| Capability | RawCull 3.2 | Leading documented pattern | Evaluation |
|---|
| Embedded RAW preview speed | Yes, with memory/disk caches | Expected in every specialist culler | Competitive |
| Developed RAW preview | Yes | Available in broader editing products | Competitive for review |
| Camera AF-point evidence | Sony and Nikon | Rarely exposed as ranking evidence | RawCull advantage |
| Explainable sharpness | Full-frame, subject, AF region, confidence, cautions | Usually a single focus/blur score | RawCull advantage |
| Visual burst grouping | Neighboring-frame similarity | Duplicate stacks, scenes, auto-stacks | Competitive |
| Side-by-side comparison | Manual and burst workspaces | Survey and synchronized close-ups | Competitive; face synchronization missing |
| Semantic text search | Local CLIP | Not central in most specialist cullers | RawCull advantage |
| Face close-ups | No dedicated surface | Prominent in Narrative, Aftershoot, Lightroom | Major gap |
| Eye-open assessment | No | Standard for people photography | Major gap |
| Whole-shoot automatic first pass | No | Automated Keep/Review/Reject or equivalent | Major gap |
| Target number/percentage | No | Imagen; staged/advertised elsewhere | Gap |
| Exposure and misfire rejection | No automatic reject category | Lightroom and Imagen | Gap |
| Genre-aware moment selection | No | FilterPixel and other automated cullers | Long-term gap |
| Learning from corrections | No | Advertised by Aftershoot and FilterPixel | Long-term gap |
| Stars/reject state | Internal persistence | Standard | Present but isolated |
| Color labels and flags | No standard interchange | Common professional workflow | Gap |
| XMP read/write | No | Essential for Lightroom/Capture One handoff | Highest workflow gap |
| RAW formats | ARW, NEF, and DNG | Major RAW families plus JPEG/HEIC | Large addressable-market gap remains |
| Safe card ingest | Selected-folder workflow | Rename, metadata, backup, verification | Partial |
| Local/offline inference | Yes | Also supported by some competitors | Strong, not unique alone |
| Exact duplicate files across folders | No | Lightroom catalog function | Low-priority gap |
Recommended Roadmap
P0: Maintain Release Qualification For 3.2.2
Do not add a large new culling feature during the final release window. Finish
the contract already presented to users:
- Run the complete automatic and manual acceptance matrix on representative
ARW, NEF, and DNG catalogs.
- Verify clean DataComp CLIP and SAM 3 installs, cancellation, removal,
relaunch, invalid-model recovery, explicit SAM licence acceptance, and Vision
fallback on the release candidate.
- Exercise Deep Review multi-subject masks and cached subject outlines across
loupe, zoom, comparison, and the review sheet.
- Run model-provenance and release-metadata gates against the published
v3
archives and manifest. - Record performance and peak memory for small, medium, and large catalogs.
- Preserve a non-AI path whose basic culling workflow does not depend on model
availability.
The current production code points to the GitHub v3 model manifest and
filters the catalog to DataComp CLIP and Meta SAM 3. Both descriptors are
ready; SAM 3 requires acceptance of the checksum-verified bundled licence.
OpenAI CLIP and EfficientSAM remain prepared but excluded. Release validation
is implemented in make verify-model-provenance and make release-preflight.
P1: XMP interoperability
This remains a strong candidate for the first substantial interoperability
feature after 3.2.2.
Required behavior
- Read existing XMP sidecars on catalog load.
- Import standard stars, reject flags, pick state, and color labels.
- Optionally import ratings written in-camera.
- Write RawCull decisions to a sidecar without modifying RAW bytes.
- Preserve unrelated XMP namespaces and fields.
- Detect whether the XMP changed since it was read.
- Present a conflict decision instead of overwriting newer metadata.
- Support batch write, cancellation, and atomic replacement.
- Offer explicit Lightroom Classic and Capture One label mappings.
- Provide Write XMP, Reveal in Finder, and Open in editing app actions.
Persistence rule
RawCull’s internal JSON should remain the durable application record. XMP is an
interchange projection. That permits RawCull-specific analysis and review state
to evolve without placing private schemas into sidecars.
flowchart LR
RAW["RAW source"] --> READ["Read existing XMP"]
READ --> MERGE["Merge into RawCull rating state"]
UI["User culling decisions"] --> STATE["RawCull JSON"]
STATE --> PROJECT["Project standard stars, flags, and labels"]
PROJECT --> CONFLICT{"Sidecar unchanged?"}
CONFLICT -->|yes| WRITE["Atomic XMP write"]
CONFLICT -->|no| REVIEW["Show conflict and preserve both choices"]Acceptance criteria
- Round-trip a representative sidecar through Lightroom Classic.
- Round-trip stars and compatible labels through Capture One.
- Prove unrelated IPTC and application-specific fields are unchanged.
- Handle read-only catalogs, missing sidecars, malformed XML, and concurrent
external edits.
- Never modify the RAW source while exporting rating metadata.
P1: Face and eye inspection
People-specific assessment is the clearest feature gap against Narrative,
Aftershoot, and Lightroom.
Proposed interface
- A face strip beside the loupe and burst comparison.
- One crop for every important face, not only the largest face.
- Synchronized crops across the frames in a burst.
- Per-face sharpness and eye-region sharpness.
- Eye state: Open, Closed, Uncertain, or Not visible.
- Occlusion, profile, small-face, and low-resolution cautions.
- Main-subject versus background-face classification.
- An Everyone acceptable summary for group photographs.
- A shortcut that moves directly between questionable faces.
Eye state must not be reduced to a forced binary value. Sunglasses, profile
faces, motion, hair, tiny background faces, and deliberate expression require
an uncertain state. Automatic first-pass policy should send uncertainty to
Review rather than Likely Reject.
Evidence model
Each face result should record:
- normalized bounding box;
- face identity only within the active catalog or burst, unless the user
explicitly enables persistent people grouping;
- face sharpness and eye-region sharpness;
- eye-state result and confidence;
- source fingerprint and analysis descriptor;
- whether the face is considered a main subject; and
- the reasons shown in the UI.
There is no dedicated face/eye-quality model in Apple’s current public Core AI
model catalog. This feature therefore needs Apple Vision face observations,
possibly a purpose-trained eye-state model converted to Core AI, or both. A
general object detector or VLM should not be presented as a reliable substitute
without a labeled validation set.
P1: Explainable First Pass
RawCull already produces much of the technical evidence needed for a useful
first pass. The first implementation can be a deterministic policy engine; it
does not need another neural network.
Output queues
| Queue | Meaning | Examples |
|---|
| Keep | Strong candidate with sufficient evidence | Best in burst, sharp subject, acceptable faces |
| Review | Creative or uncertain decision | Slight motion, uncertain eye state, disagreement between AF and subject evidence |
| Likely Reject | Strong technical reason to de-prioritize | Clearly missed focus, inferior duplicate, confirmed closed eyes in a posed group |
Safety invariants
- Never delete a source photograph.
- Always keep access to every queue.
- Retain at least one candidate per burst unless the user explicitly changes
the rule.
- Put low-confidence or conflicting evidence in Review.
- Explain every automatic placement using stored evidence.
- Permit one-click override and undo.
- Re-running with changed policy must not erase manual ratings or winners.
- Persist policy identity so old results are not misrepresented after an
algorithm change.
Initial evidence
- burst membership and within-burst rank;
- visual similarity distance;
- full-frame, salient-subject, and AF-region sharpness;
- AF-point containment and distance from the salient subject;
- analysis confidence and existing caution reasons;
- exposure clipping derived from image data rather than histogram appearance
alone;
- face and eye results when available; and
- existing manual ratings and winners as hard overrides.
Target yield
After Keep/Review/Reject is trustworthy, add a requested target count or
percentage. The target is a policy constraint, not evidence of quality. RawCull
should show when reaching the requested number requires inclusion of weaker or
lower-confidence frames.
P2: Local preference profiles
RawCull can learn useful preferences by adjusting transparent policy weights
before attempting end-to-end personalized AI.
Possible learned values include:
- preferred number of retained frames per burst;
- tolerance for subject motion and deliberate blur;
- relative importance of salient-subject sharpness and camera AF evidence;
- preferred subject size and framing;
- desired delivery percentage;
- treatment of uncertain eyes and background faces; and
- preference for technical perfection versus moment uniqueness.
Profiles should be named, inspectable, exportable, resettable, and scoped by
genre when the user wants that behavior. A profile must never silently turn a
manual reject into a keeper or replace a manual burst winner.
Suggested first profiles:
- Portrait and group;
- Wedding and event;
- Sports and action;
- Wildlife and birds;
- Landscape and architecture; and
- General/manual assistance.
Genre names should initially select documented weights and rules. Claims such
as emotion, storytelling, or peak action require separate labeled
evaluation before they become user-visible facts.
The market supports more camera systems than RawCull. A practical expansion
order is:
- Canon CR3;
- Fujifilm RAF;
- DNG;
- JPEG and HEIC/HEIF companions;
- Panasonic RW2, Olympus ORF, and Pentax PEF.
Generic culling does not have to wait for MakerNote AF support. A format can
provide embedded preview, EXIF, sharpness, similarity, ratings, and XMP while
reporting that camera AF metadata is unavailable.
The UI should distinguish three states:
- camera AF point available and used;
- camera AF point absent or unsupported, with other evidence available; and
- image decoding or analysis unavailable.
This avoids treating broader format support as inferior or silently assigning
a false AF coordinate.
P2: Safer ingest
RawCull should adopt a focused subset of Photo Mechanic’s ingest strengths:
- camera-card detection;
- primary and backup destinations;
- optional checksum verification;
- configurable folder and filename templates;
- preservation of in-camera ratings;
- copy progress and cancellation;
- an ingest receipt containing source, destinations, counts, and failures; and
- optional eject only after successful verification.
Do not erase or format source media. Rich IPTC templates and sports code
replacement can remain outside RawCull unless users demonstrate demand.
P3: Exact duplicates and catalog maintenance
Exact byte duplicates across folders are different from visually similar burst
frames. They require content hashing, catalog scope, and careful deletion or
move policy. This is useful but is closer to digital-asset management than
culling. It should follow the workflow features above.
AI Evaluation
Principles for adding another model
The existence of an Apple export recipe does not mean that RawCull should ship
the model. Every additional bundle creates:
- a download and storage cost;
- first-use specialization time;
- memory and energy pressure;
- another license and redistribution decision;
- another model fingerprint and cache-compatibility dimension;
- failure, cancellation, update, and removal states;
- a benchmark and regression obligation; and
- user-interface complexity.
A new model is justified only when it solves a measured culling problem better
than current Vision, CLIP, segmentation, deterministic analysis, or a small
purpose-trained model.
Apple’s Core AI Models repository
provides export recipes and Swift runtime utilities for macOS and iOS 27. The
current upstream catalog
includes language models, diffusion models, Qwen3-VL, CLIP, Depth Anything v3,
EDSR, EfficientSAM, PVT v2, SAM 3, YOLOS, audio models, and text encoders.
RawCull pins an exact repository revision, so features present on upstream
main are not automatically present in the version used by a release.
Candidate model matrix
| Model or family | Possible RawCull use | Product fit | Recommendation |
|---|
| EfficientSAM ViT-Tiny | Subject masks from points, boxes, or a point grid; subject-detail focus scoring | High; existing PhotoAIKit and RawCull backend | Evaluate first and aim to make available after release gates |
| SAM 3 | Text-guided subject segmentation and Deep Review | High capability, but 848M parameters and gated redistribution | Keep optional; do not block First Pass on it |
| CLIP ViT-B/32 | Existing semantic search, similarity, zero-shot labels | Already central | Stabilize and benchmark; avoid model proliferation |
| YOLOS Tiny/Base | Object boxes, subject occupancy, possible ball/animal/person evidence | Medium; fixed object vocabulary and not a quality model | Prototype Tiny only after face/eye and First Pass |
| Qwen3-VL 2B | Local captions, scene summaries, possible moment/composition suggestions | Interesting but costly and generative | Research preview only; never technical ground truth |
| Depth Anything v3 Small | Foreground/background separation, depth layering, background-distraction evidence | Medium-low; segmentation and saliency already overlap | Experiment only if a benchmark proves incremental value |
| PVT v2 B0 | Small visual backbone for a custom classifier | Low by itself; no culling-quality head | Use only as a base for a purpose-trained model |
| EDSR x2 | Sharper-looking display crop | Poor for culling evidence because reconstruction can invent detail | Do not use for scoring; optional preview only |
| RoBERTa/T5 | Query normalization or structured text processing | Low; CLIP queries and deterministic UI do not require them | Do not ship for current workflows |
| Qwen/Gemma/Mistral/GPT-OSS LLMs | Natural-language explanations from structured evidence | Low relative to size and complexity | Prefer deterministic explanations; no near-term bundle |
| Stable Diffusion/FLUX | Image generation | No culling purpose and risks changing evidence | Exclude |
| CLAP/Whisper/Wav2Vec | Audio understanding/transcription | No current still-photo culling purpose | Exclude |
1. EfficientSAM: strongest next candidate
Apple’s current EfficientSAM recipe describes a 10-million-parameter ViT-Tiny
model. It supports a foreground click, box prompt, multiple point queries, and
a segment-everything point grid. RawCull already has:
- a
CoreAIEfficientSAMBackend dependency; - an EfficientSAM provider path in
RawCullAIIntegration; - an EfficientSAM model directory;
- a
RawCullSegmentationModel.efficientSAM identity; and - Deep Review code capable of consuming subject masks.
The missing work is primarily product inclusion, model packaging, validation,
license/provenance clearance, download-catalog support, and quality evaluation.
Recommended use:
- make EfficientSAM the lightweight Deep Review candidate;
- discover likely foreground subjects with a bounded point grid;
- combine masks with saliency and camera AF points;
- let the user click or box the intended subject when automatic discovery is
ambiguous; and
- cache masks using the full model and prompt identity.
Required evaluation:
- portraits, groups, animals, birds, sports, landscapes, and low-light scenes;
- subject selection success rather than generic segmentation metrics alone;
- mask placement and boundary quality;
- AF-point containment stability;
- detail-score stability inside the selected mask;
- false selection of background objects;
- latency, peak memory, first-use specialization, cache size, and energy; and
- quality comparison against SAM 3 on the same labeled cases.
EfficientSAM does not understand a text target. The UI must describe point-grid
or user-prompted selection accurately and must not imply SAM 3-style language
grounding.
2. SAM 3: premium text-guided Deep Review
SAM 3 remains valuable when a photographer wants to specify bird, face,
player, or another text target. Apple’s recipe describes an 848-million-
parameter gated model. The capability is stronger than EfficientSAM’s point
prompts, but the size, gated access, and redistribution review make it a poor
mandatory dependency.
Recommended policy:
- keep Vision and ordinary burst review fully functional without it;
- treat SAM 3 as an optional advanced model;
- make its managed download visible only after legal and provenance clearance;
- require explicit license acceptance if the final legal review requires it;
- compare its added winner-selection value with EfficientSAM, not only mask
appearance; and
- avoid advertising Deep Review as turnkey while normal users cannot obtain
the model through the app.
3. Dedicated face and eye model: needed but not in the catalog
The most important new AI capability is face/eye quality, yet Apple’s public
Core AI model catalog currently contains no dedicated model for:
- eyes open versus closed;
- eye-region sharpness;
- facial-expression suitability;
- gaze or camera engagement;
- group-photo all-faces acceptance; or
- intentional versus accidental eye closure.
RawCull should first evaluate Vision face rectangles and landmarks combined
with the existing sharpness analyzer. If eye-state classification remains
insufficient, a small purpose-trained classifier is a better fit than adding a
general 2B vision-language model.
Such a classifier needs consented or properly licensed training/evaluation
data, coverage across skin tones and ages, glasses and sunglasses, profiles,
occlusions, makeup, low light, motion, and small faces, plus explicit
uncertainty calibration. Accuracy must be reported per condition, not only as a
single aggregate percentage.
4. YOLOS: possible object evidence
Apple provides YOLOS Tiny at approximately 6.5 million parameters and YOLOS
Base at approximately 127 million parameters, together with an object-detection
runtime product. A detector could provide:
- person, animal, vehicle, or sports-object boxes;
- subject occupancy and edge-cutoff cautions;
- possible ball-in-frame evidence for selected sports; and
- a subject ROI when saliency is ambiguous.
Limitations:
- a general object vocabulary does not equal photographic importance;
- detection confidence does not measure focus, expression, composition, or
peak action;
- it cannot replace a dedicated face/eye assessment; and
- unsupported subjects could bias the first-pass policy.
If evaluated, begin with YOLOS Tiny and use detections as optional evidence,
never as an automatic reject condition.
5. Qwen3-VL: later research only
Apple’s current upstream main contains a Core AI export path for
Qwen3-VL-2B-Instruct with a 448-pixel vision encoder, token embedding, text
decoder, tokenizer, and bundle metadata. This capability may be newer than the
exact coreai-models revision pinned by RawCull.
Potential experiments:
- create searchable catalog captions;
- summarize a burst’s visible differences;
- suggest scene or genre labels;
- identify possible emotional or peak-action moments; and
- turn structured evidence into accessible natural language.
Risks:
- generative claims can hallucinate details or intent;
- a 448-pixel view may miss the technical detail used for focus decisions;
- per-image inference across a large catalog may be too slow or memory-heavy;
- captions add a new private persistent-data category;
- a VLM can sound more certain than its evidence; and
- the model and tokenizer create a much larger distribution obligation.
Any VLM result must be labeled as a suggestion. It must not override measured
sharpness, face results, user ratings, or manual winners. A small deterministic
formatter remains preferable for explanations such as “subject sharpness was
higher and the AF point was inside the selected mask.”
6. Depth Anything v3: limited incremental value
Depth Anything v3 Small predicts monocular depth and confidence. It might help
measure subject separation, background complexity, foreground obstructions, or
depth-layer composition.
However, RawCull already has saliency and optional segmentation. Depth is not
itself a culling-quality measure, and incorrect monocular depth can create
confident but irrelevant rankings. Evaluate it only against a labeled feature
such as “background distraction” or “subject separation.” Do not add it merely
because a Core AI recipe exists.
7. EDSR: never use reconstructed detail for sharpness scoring
EDSR can enlarge a low-resolution preview, but super-resolution creates an
estimate of detail rather than evidence from the source. It must not feed:
- sharpness scoring;
- AF-region scoring;
- focus masks;
- eye sharpness;
- burst winner selection; or
- any label presented as source-image quality.
An optional display-only enhancement could be considered, but RawCull can
already extract embedded previews or develop the RAW. That makes EDSR a low
priority even for presentation.
Model Admission Gates
Every new model should pass all gates below before it appears in production
Settings.
Product gate
- Name the user problem and the decision improved by the model.
- Define a non-model baseline.
- Demonstrate incremental value on representative RawCull catalogs.
- Confirm that the feature remains understandable and reversible.
Quality gate
- Use a versioned, labeled evaluation set.
- Separate technical metrics from subjective photographer preference.
- Record false-positive and false-negative costs.
- Calibrate an uncertain state where applicable.
- Compare results across supported cameras, genres, and difficult conditions.
- Require human review of ranking changes, not only tensor parity.
Conversion and identity gate
- Pin the exact upstream revision and weight checksum.
- Record the exact Core AI exporter and PhotoAIKit revisions.
- Verify reference/Core AI parity using the real preprocessing path.
- Fingerprint the complete runtime bundle and tokenizer/resources.
- Include preprocessing, normalization, dimensions, and output interpretation
in the backend descriptor.
- Invalidate only incompatible cached artifacts after a model change.
- Measure cold specialization and warm inference separately.
- Record median and tail latency per image.
- Record peak resident memory and memory-pressure behavior.
- Test bounded concurrency and cancellation.
- Measure a realistic 500-, 2,000-, and 10,000-image workflow where relevant.
- Confirm that UI interaction and thumbnail loading remain responsive.
Distribution gate
- Complete license and redistribution review for weights and converted assets.
- Preserve model card, notices, upstream revision, export command, and hashes.
- Verify the archive produced by Managed Background Assets.
- Test download, cancellation, validation, relaunch, update, removal, and
insufficient-storage behavior.
- Do not confuse technical success with permission to redistribute.
Presentation gate
- State whether a result is measured, inferred, or generated.
- Show confidence and meaningful uncertainty.
- Do not hide unavailable models behind unexplained disabled controls.
- Provide fallback behavior and recovery steps.
- Maintain VoiceOver descriptions for progress, result, confidence, and error
states.
Proposed AI Model Roadmap
| Stage | Models | Goal |
|---|
| 3.2 release | Vision plus DataComp/OpenAI CLIP | Stable similarity, burst grouping, and semantic search |
| 3.3 candidate | EfficientSAM | Lightweight, obtainable Deep Review and user-prompted subject masks |
| Later optional | SAM 3 | Text-guided premium Deep Review after redistribution clearance |
| Parallel research | Vision plus a small eye-state classifier | Face/eye inspection and group-photo review |
| Later experiment | YOLOS Tiny | Optional object/subject evidence |
| Research only | Qwen3-VL 2B | Captions and subjective suggestions, never technical ground truth |
| Evidence-dependent | Depth Anything v3 | Background/subject-separation evidence only if benchmarks justify it |
This ordering solves product needs rather than maximizing the number of models
shown in Settings.
Features Not Recommended
RawCull should defer or reject the following unless its product scope changes:
- RAW editing profiles and automatic editing;
- retouching and generative removal;
- cloud galleries, proofing, print sales, and delivery;
- diffusion-based image generation;
- destructive automatic rejection or deletion;
- an LLM used only to rewrite deterministic evidence in friendlier words;
- super-resolution used as quality evidence;
- persistent face recognition by default; and
- a large model marketplace without a specific culling purpose for every
model.
Aftershoot and Imagen are developing broad cull-edit-retouch-deliver systems.
Competing on their complete scope would consume resources without strengthening
RawCull’s most defensible advantages.
Success Measures
Future features should be judged by user outcomes rather than by model count.
Useful measures include:
- time from catalog open to completed first pass;
- time spent zooming into faces;
- percentage of automatic placements the photographer changes;
- missed-keeper rate in Likely Reject;
- number of unresolved Review images;
- agreement with manually chosen burst winners;
- XMP round-trip success rate;
- catalog formats and cameras admitted successfully;
- peak memory and cancellation latency; and
- percentage of workflows completed without leaving RawCull before handoff.
For automatic culling, missed keepers matter more than a superficially high
overall accuracy score. The default policy should be conservative enough that
uncertain photographs remain visible in Review.
Recommended Next Design Work
Before implementation, prepare three focused specifications:
- an XMP mapping and conflict-resolution specification;
- a face/eye evidence schema and labeled acceptance matrix; and
- a First Pass policy document defining queue invariants, reasons, target
yield, overrides, persistence identity, and failure behavior.
EfficientSAM can be evaluated in parallel because its backend boundary already
exists. It should not delay XMP or the deterministic First Pass, and a new model
should never be required merely to display existing RawCull evidence clearly.
10 - Artificial Intelligence
RawCull AI architecture, model downloads, and Objects test-release status.
Artificial Intelligence in RawCull
This section explains how RawCull uses reusable AI components without allowing
model-runtime details to spread through the application. It is written as a
learning path: first understand the boundary between RawCull and PhotoAIKit,
then study the package itself, and finally follow CLIP from application startup
to a persisted similarity artifact.
The source for this section comes from the pinned PhotoAIKit dependency and the
RawCull repository. Paths beginning with Sources/ refer to PhotoAIKit at the
revision in Package.resolved. RawCull intelligence code lives under
RawCull/Intelligence; app composition and presentation remain under
RawCull/Main, RawCull/Model, and RawCull/Views.
What This Section Documents
| Document | Main question | Start here when |
|---|
| This overview | Where does AI belong in the system? | You need the vocabulary and responsibility split |
| The RawCull AI Runtime | How are providers and long-lived features assembled? | You are tracing startup, refresh, or service replacement |
| AI Models in RawCull | How do CLIP, Vision, SAM 3, Qwen, and Objects work in the app? | You are tracing an analysis from input to result |
| Download AI models | How are the three release packs rebuilt? | You are preparing source weights, conversions, or archives |
| AI Model Licence and Provenance Clearance | What evidence is required before a model can ship? | You are reviewing licences, provenance, or release readiness |
| Publishing and Testing RawCull AI Models | How are packs published and Objects tested? | You are preparing a model release or updating the download manifest |
PhotoAIKit and RawCull use these AI backends:
- CLIP image embeddings for visual similarity and semantic search.
- SAM 3 and EfficientSAM subject segmentation for subject masks.
- Apple Vision feature prints for always-available image similarity.
- Qwen3-VL for local photo analysis and for concept discovery and assessment
of numbered SAM 3 objects.
The 3.2.6 production download catalog exposes DataComp CLIP, Meta SAM 3,
and Qwen3-VL-2B-Instruct. OpenAI CLIP remains excluded and EfficientSAM is not
a production download. SAM 3 requires acceptance of its verified bundled
licence before download. The App Store build uses Apple-hosted Managed
Background Assets; the Direct/Developer ID build retains a self-hosted v3
manifest. See AI Model Downloads for the exact distinction.
Vision is the startup and service-selection fallback: RawCull uses it when CLIP
is disabled or the selected CLIP bundle cannot produce a validated provider. A
selected CLIP indexing pass keeps its valid per-file artifacts and records the
files that fail; it does not mix Vision artifacts into that pass or
automatically rerun the whole batch.
The AI Models in RawCull guide follows code connected
to similarity, semantic search, burst analysis, Deep Review, Qwen, and Objects.
The PhotoAIKit architecture guide covers SAM 3
contracts, workflows, and storage so that the package design is understandable. Not every reusable package capability is necessarily
exposed as a finished RawCull user workflow.
Objects test-release status
The maintainer is preparing a user test of AI Analysis → Objects. The
feature combines Qwen concept discovery, separate SAM 3 masks, and Qwen
assessment of numbered crops. The structural decoder checks board IDs and
response fields; it cannot verify that generated text matches the photograph.
A recent two-puffin result had two retained objects but an invented third bird
in its summary, and a separate two-puffin result has an unresolved
crop/description mismatch. Treat Objects output as advisory and verify it
against the source. See Publishing and Testing RawCull AI Models
for the known issues, tester checks, and release evidence still needed.
The Central Design Idea
RawCull and PhotoAIKit answer different kinds of questions.
PhotoAIKit asks:
- What does an image source look like at a package boundary?
- How is a model bundle validated and identified?
- How does a backend produce and compare a similarity artifact?
- How is bounded indexing or segmentation orchestrated?
- How can reusable artifacts and masks be encoded or cached?
RawCull asks:
- Where are models installed for this application?
- How is a Sony, Nikon, or DNG RAW file decoded for AI input?
- Which backend did the user request?
- When should a catalog be indexed or reindexed?
- How do similarity distances affect burst grouping and culling?
- What state and wording should SwiftUI present?
This is dependency inversion in practical form. PhotoAIKit defines small
protocols such as ImageDecoding, ImageSimilarityArtifactProviding, and
ImageSimilarityArtifactComparing. RawCull injects app-specific implementations
and keeps its file model, UI, sandbox policy, and culling decisions outside the
package.
flowchart LR
UI["RawCull SwiftUI and settings"] --> Integration["RawCullApplicationState composition root"]
Integration --> AppAdapter["RawCull adapters: paths, RAW decoding, policy"]
AppAdapter --> Contracts["PhotoAIContracts"]
Integration --> CLIP["CoreAICLIPBackend"]
Integration --> SAM3["CoreAISAM3Backend"]
Integration --> Vision["VisionFeaturePrintBackend"]
AppAdapter --> Workflows["PhotoAIWorkflows"]
Workflows --> Contracts
Storage["PhotoAIStorage"] --> Contracts
CLIP --> Contracts
SAM3 --> Contracts
Vision --> ContractsThe arrow direction matters: PhotoAIKit does not import RawCull. A reusable
package should not need to know what FileItem, RawCullViewModel, an
app-specific model folder, or a burst winner means.
Responsibility Boundary
| Concern | Owner | Reason |
|---|
| Typed model, image, artifact, and segmentation contracts | PhotoAIKit | Backends and hosts need one stable language |
| Core AI CLIP and SAM 3 inference | PhotoAIKit backend products | Framework-specific tensor and inference code is reusable |
| Vision feature-print generation and native distance | PhotoAIKit backend product | The opaque Vision payload stays behind its backend boundary |
| Bounded indexing, optional fallback mechanisms, segmentation, and mask selection | PhotoAIKit workflows | These mechanisms do not depend on RawCull UI or culling policy |
| Optional embedding codecs and mask stores | PhotoAIKit storage | Persistence mechanics are reusable, but locations are not |
| Model installation directories and candidate order | RawCull | Paths and sandbox policy belong to the host application |
| RAW decoding | RawCull | PhotoAIKit should not depend on RawParserKit or camera formats |
| Settings and capability wording | RawCull | User-facing state and localization belong to the app |
| Similarity ranking adjustments and burst grouping | RawCull | These are photo-culling product decisions, not CLIP behavior |
| Burst-analysis cache location and lifecycle | RawCull | The host owns when and where catalog results persist |
Current Runtime Shape
RawCullApplicationState is the object-graph assembly boundary. It creates one
RawCullAIModelRuntime, one shared SimilarityScoringModel, the focused
similarity and semantic-search features, Deep Review, Qwen and Objects features,
the main view model, and one RawCullIntelligenceRuntime. Identity assertions
protect against accidentally constructing parallel observable state.
RawCullAIModelRuntime owns concrete providers, resource managers, Qwen
inference, and separate subject/object mask stores. RawCullIntelligenceRuntime
owns stable feature lifetimes and applies complete revisioned configurations
from settings. See The RawCull AI Runtime for the full
construction and refresh sequence. Views receive focused feature surfaces
instead of the composition root or low-level scoring model:
| Consumer | Narrow dependency | Capability and persistence rule |
|---|
RawCullSimilarityFeature | Shared SimilarityScoringModel plus RawCullSimilarityServicing | Owns the public hydration, indexing, ranking, cancellation, and backend-presentation surface while persisting descriptor-valid artifacts |
RawCullSemanticSearchFeature | Shared scoring model and optional semantic-search service | Projects semantic-search state, binds weakly to application selection/navigation, and never indexes missing images as a query side effect |
BurstAnalysisCoordinator | Similarity feature, scoring models, and cache repository | Owns burst generation, progress, cache preparation, missing computation, grouping, ranking, cancellation, and derived-cache saving |
DeepAIReviewController | DeepAIReviewFeature | Builds immutable requests from app evidence and validates the group signature before recommendations reach culling policy |
RawCullAISettingsModel | RawCullIntelligenceConfigurationApplying | Publishes one ordered configuration; the runtime ignores stale revisions and applies only meaningful identity changes |
The safe startup and refresh path is:
RawCullApp calls RawCullApplicationState.live() and retains its view
model and intelligence runtime as stable @State roots.- Assembly creates the shared scoring model and focused features from the
initial Vision-backed configuration.
RawCullAISettingsModel.refresh() asks the model runtime to validate both CLIP
and both segmentation-model candidates.- PhotoAIKit validates model bundles and derives model-asset fingerprints.
- Settings publishes a monotonically revisioned configuration. The runtime
replaces similarity or semantic-search services only when their identities
changed and applies segmentation selection independently.
- Missing, invalid, or disabled CLIP leaves burst similarity on Vision and
semantic search unavailable.
- CLIP indexing retains valid files and logs per-file failures; Vision is not
inserted into that CLIP result set.
- Changing the segmentation selection immediately updates the active provider.
A full refresh rechecks every candidate and saved-evidence state.
stateDiagram-v2
[*] --> VisionStartup
VisionStartup --> CheckingModels: settings refresh
CheckingModels --> Configuration: publish newer configuration revision
Configuration --> CLIPSelected: preference enabled and provider ready
Configuration --> VisionSelected: selected CLIP unavailable
CLIPSelected --> CLIPArtifacts: keep valid per-file artifacts
CLIPSelected --> PartialCLIP: record and exclude failed files
VisionSelected --> VisionArtifacts: index catalog
CheckingModels --> SegmentationSelected: activate SAM 3 or EfficientSAMBurst similarity, semantic search, and Deep Review are separate features even
when they share package code. Burst similarity may use Vision or CLIP. Semantic
search requires CLIP image artifacts whose descriptor exactly matches the text
provider. Deep Review uses segmentation masks and has its own availability,
selection, and storage lifecycle.
Vocabulary
| Term | Meaning in this codebase |
|---|
| Provider | A backend object that performs inference or creates an artifact |
| Backend descriptor | Identity of the backend, model, representation, preprocessing, normalization, and configuration |
| Similarity artifact | A descriptor plus a backend-owned payload; CLIP stores an encoded vector, while Vision stores an opaque archived observation |
| Source fingerprint | Standardized file path, size, and modification date used to detect changed source images |
| Model fingerprint | Identity derived from the selected .aimodel or .aimodelc, cryptographically verified when the manifest provides a checksum |
| Composition root | The one place where concrete providers, stores, paths, and app adapters are assembled |
| Partial CLIP result | Valid CLIP artifacts plus per-file failures; failed files remain unavailable to similarity and burst grouping until a later successful index |
| Host | The application integrating PhotoAIKit; here, RawCull |
10.1 - Download AI models
Download and prepare the three AI model packs
This is a terminal workbook for rebuilding RawCull’s DataComp CLIP, Meta
SAM 3, and Qwen3-VL-2B-Instruct model packs from source and finishing with
three Apple Managed Background Assets .aar files. It was assembled on
September 26, 2026 from the local PhotoAIKit exporters, the RawCull release
runbook, and the existing layout in /Users/thomas/ModelAssets/Release.
Run each numbered block in zsh and inspect the indicated output before the
next block. The workbook builds in a fresh directory below
/Users/thomas/ModelAssets; it does not overwrite the three existing
Release/Output/*.aar files. Conversion is CPU and memory intensive and the
source downloads are several gigabytes. Keep ample free space for source
weights, intermediate models, three converted bundles, and three archives.
These commands produce local release candidates, not App Store approval.
After packaging, compare the new hashes with the application catalog and follow
the publishing runbook
to update RawCull, upload the packs, and test a signed TestFlight build. A fresh conversion may produce different archive
bytes even when the same source model was used.
0. What will be produced
| Model | Permanent pack ID | Selected installed model path | Output archive |
|---|
| DataComp CLIP | rawcull-clip-datacomp | Models/CLIP-DataComp | clip-datacomp.aar |
| Meta SAM 3 | rawcull-sam3 | Models/SAM3 | sam3.aar |
| Qwen3-VL-2B-Instruct | rawcull-qwen3-vl-2b | Models/Qwen/qwen3_vl_2b | qwen3-vl-2b.aar |
The starting evidence is the RawCull repository at
/Users/thomas/GitHub/RawCull/RawCull, its sibling
/Users/thomas/GitHub/RawCull/PhotoAIKit, and the notices in
RawCull/ModelAssets/Notices. The old archive sizes and SHA-256 values in the
application catalog describe the September 17, 2026 artifacts, not expected
values for this rebuild. OpenAI CLIP and EfficientSAM are outside this
three-pack release.
Install Xcode 27 with the Core AI and Background Assets tools and select it in
Xcode Settings or with xcode-select. Install git and uv if needed. The
commands below merely check the selected tools:
set -e
set -o pipefail
RAWCULL_REPO='/Users/thomas/GitHub/RawCull/RawCull'
PHOTOAIKIT_REPO='/Users/thomas/GitHub/RawCull/PhotoAIKit'
MODEL_ASSETS='/Users/thomas/ModelAssets'
test -d "$RAWCULL_REPO/.git"
test -d "$PHOTOAIKIT_REPO/.git"
test -d "$MODEL_ASSETS"
command -v uv
command -v git
command -v python3
xcode-select -p
xcodebuild -version
xcrun ba-package --version
If uv is missing and Homebrew is available, install it with brew install uv, then rerun the checks. Record the exact tool versions with the eventual
archive hashes. At the time this workbook was written the local Mac selected
Xcode 27.0 (27A266a) and ba-package 2.0; a later toolchain can produce
different output. The PhotoAIKit exporter scripts declare their Python
dependencies in their own # /// script blocks, so uv run creates the
appropriate isolated environments automatically. Do not combine the CLIP and
SAM dependency sets by hand: their transformers requirements differ.
2. Start a clean build directory
Paste this block into the same Terminal window that will run later blocks. If
you open a new window, rerun the variable definitions in this block and set
BUILD_ROOT to the directory printed by echo.
set -e
set -o pipefail
RAWCULL_REPO='/Users/thomas/GitHub/RawCull/RawCull'
PHOTOAIKIT_REPO='/Users/thomas/GitHub/RawCull/PhotoAIKit'
MODEL_ASSETS='/Users/thomas/ModelAssets'
BUILD_ROOT="$(mktemp -d "$MODEL_ASSETS/Build-2026-09-26.XXXXXX")"
HF_HOME="$BUILD_ROOT/HuggingFace"
HF_HUB_CACHE="$HF_HOME/hub"
export HF_HOME HF_HUB_CACHE
mkdir -p "$BUILD_ROOT/Release/Models/Qwen" \
"$BUILD_ROOT/Release/Notices" \
"$BUILD_ROOT/Release/Packaging" \
"$BUILD_ROOT/Release/Output" \
"$BUILD_ROOT/Release/Evidence" \
"$BUILD_ROOT/Tools"
echo "BUILD_ROOT=$BUILD_ROOT"
df -h "$MODEL_ASSETS"
This layout isolates the Hugging Face cache and the new release candidates.
Never run rm -rf against the existing ModelAssets/Release tree to make room.
For a later run, change the date in the mktemp template or leave it: the
random suffix still creates a separate directory.
3. Freeze source and exporter revisions
The model revisions recorded in the existing RawCull catalog are:
CLIP_REV='4afec35ffe57a943d569ff7ee888061830164da8'
SAM_REV='3c879f39826c281e95690f02c7821c4de09afae7'
QWEN_REV='78448d793a7eb2f7a987a1da76d464384aa1becd'
CLIP_TOKENIZER_REV='3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268'
COREAI_MODELS_REV='475c585fdb0fe82a83c8f777f259e9414bd44c98'
COREAI_MODELS_REV is the revision pinned by the local PhotoAIKit package in
the September 2026 RawCull work. These revision labels are a starting recipe.
The old DataComp and Qwen provenance explicitly does not prove which exact
weight bytes their exporters consumed. This rebuild records the actual source
inventory so its evidence is stronger.
First save the repository states. If either repository has local changes to
the exporter, review them before proceeding:
git -C "$RAWCULL_REPO" rev-parse HEAD | tee "$BUILD_ROOT/Release/Evidence/rawcull-commit.txt"
git -C "$PHOTOAIKIT_REPO" rev-parse HEAD | tee "$BUILD_ROOT/Release/Evidence/photoaikit-commit.txt"
git -C "$RAWCULL_REPO" status --short
git -C "$PHOTOAIKIT_REPO" status --short
shasum -a 256 "$PHOTOAIKIT_REPO/Tools/export_clip.py" \
"$PHOTOAIKIT_REPO/Tools/export_sam3.py" \
"$PHOTOAIKIT_REPO/Tools/select_sam3_asset.py" \
| tee "$BUILD_ROOT/Release/Evidence/photoaikit-exporters-sha256.txt"
Get Apple’s converter at the selected revision in the build directory. A
normal git clone also preserves its licence and Python project configuration.
If the pinned revision is unavailable, stop and choose a reviewed revision;
do not silently use the current main branch.
git clone https://github.com/apple/coreai-models.git "$BUILD_ROOT/Tools/coreai-models"
git -C "$BUILD_ROOT/Tools/coreai-models" checkout --detach "$COREAI_MODELS_REV"
git -C "$BUILD_ROOT/Tools/coreai-models" rev-parse HEAD \
| tee "$BUILD_ROOT/Release/Evidence/coreai-models-commit.txt"
4. Download immutable Hugging Face snapshots
SAM 3 is gated: sign in to Hugging Face in a browser, request/receive access to
facebook/sam3, accept its terms, and authenticate the terminal with
uvx --from huggingface-hub hf auth login. Do not paste your access token into
this workbook or a shell command. Qwen and DataComp are public, but checking
their model cards and current redistribution terms is still part of release
review. These downloads use Hugging Face’s CLI.
uvx --from huggingface-hub hf auth whoami
uvx --from huggingface-hub hf download \
laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K \
--revision "$CLIP_REV" \
--local-dir "$BUILD_ROOT/Source/CLIP-DataComp"
uvx --from huggingface-hub hf download facebook/sam3 \
--revision "$SAM_REV" \
--local-dir "$BUILD_ROOT/Source/SAM3"
uvx --from huggingface-hub hf download Qwen/Qwen3-VL-2B-Instruct \
--revision "$QWEN_REV" \
--local-dir "$BUILD_ROOT/Source/Qwen"
The exporters do not all take a --revision or --source-dir flag. To make
their ordinary model-ID lookups use the selected snapshots, populate the
isolated Hugging Face cache at the same revisions and bind its main refs to
those revisions. This is explicit, local to BUILD_ROOT, and leaves your
ordinary Hugging Face cache alone. The CLIP exporter also obtains the OpenAI
CLIP tokenizer, so cache it separately.
uvx --from huggingface-hub hf download \
laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K \
--revision "$CLIP_REV"
uvx --from huggingface-hub hf download facebook/sam3 --revision "$SAM_REV"
uvx --from huggingface-hub hf download Qwen/Qwen3-VL-2B-Instruct \
--revision "$QWEN_REV"
uvx --from huggingface-hub hf download openai/clip-vit-base-patch32 \
--revision "$CLIP_TOKENIZER_REV"
mkdir -p "$HF_HUB_CACHE/models--laion--CLIP-ViT-B-32-256x256-DataComp-s34B-b86K/refs" \
"$HF_HUB_CACHE/models--facebook--sam3/refs" \
"$HF_HUB_CACHE/models--Qwen--Qwen3-VL-2B-Instruct/refs" \
"$HF_HUB_CACHE/models--openai--clip-vit-base-patch32/refs"
printf '%s' "$CLIP_REV" > "$HF_HUB_CACHE/models--laion--CLIP-ViT-B-32-256x256-DataComp-s34B-b86K/refs/main"
printf '%s' "$SAM_REV" > "$HF_HUB_CACHE/models--facebook--sam3/refs/main"
printf '%s' "$QWEN_REV" > "$HF_HUB_CACHE/models--Qwen--Qwen3-VL-2B-Instruct/refs/main"
printf '%s' "$CLIP_TOKENIZER_REV" > "$HF_HUB_CACHE/models--openai--clip-vit-base-patch32/refs/main"
The first --local-dir downloads are the inspectable evidence copies; the
second downloads populate the cache actually used by the model-ID loaders.
Record source hashes, including every Qwen weight shard:
find "$BUILD_ROOT/Source" -type f ! -name '.DS_Store' -print0 \
| xargs -0 shasum -a 256 \
| sort > "$BUILD_ROOT/Release/Evidence/source-files-sha256.txt"
rg 'safetensors|tokenizer.json|LICENSE' \
"$BUILD_ROOT/Release/Evidence/source-files-sha256.txt" | head -80
Check that the expected source files exist before converting:
test -f "$BUILD_ROOT/Source/SAM3/model.safetensors"
test -f "$BUILD_ROOT/Source/Qwen/config.json"
find "$BUILD_ROOT/Source/Qwen" -maxdepth 1 -name '*.safetensors' -print
find "$BUILD_ROOT/Source/CLIP-DataComp" -maxdepth 1 -type f -print
DataComp binding limitation: open_clip.create_model_and_transforms uses
the datacomp_s34b_b86k preset and may resolve its checkpoint through an
OpenCLIP configuration rather than the checked-out LAION directory. Run it
with the isolated cache and inspect the export log and cache entries. Do not
claim that the LAION file was consumed solely because it was downloaded. If
the resolved source differs, capture that actual source file, revision and
SHA-256, or adapt the exporter to accept an explicit local weight path before
release.
5. Convert DataComp CLIP
The PhotoAIKit exporter creates one two-function .aimodel plus tokenizer and
bundle metadata. Its default DataComp options use ViT-B/32 at 256 px,
datacomp_s34b_b86k, float16, and static shapes. The script itself pins the
Python packages, including coreai-core==1.0.0b2 and
coreai-torch==0.4.1.
cd "$PHOTOAIKIT_REPO"
export HF_HUB_OFFLINE=1
uv run Tools/export_clip.py \
--model openclip-datacomp \
--architecture ViT-B-32-256 \
--pretrained datacomp_s34b_b86k \
--dtype float16 \
--output-dir "$BUILD_ROOT/Release/Models" \
--bundle-name CLIP-DataComp \
2>&1 | tee "$BUILD_ROOT/Release/Evidence/clip-export.log"
The exporter checks parity between the OpenCLIP tokenizer IDs and the saved
PhotoAIKit tokenizer. Stop if that check fails. HF_HUB_OFFLINE=1 intentionally
prevents a network fallback to an unpinned source. If OpenCLIP needs a different
checkpoint repository, find its actual configured repository, pin and download
it, then rerun this block in a new build directory. Do not use --overwrite
until you have saved and reviewed the first result.
test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/metadata.json"
test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/tokenizer/tokenizer.json"
test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/ViT-B-32-256-datacomp_s34b_b86k_float16_static.aimodel/main.mlirb"
cat "$BUILD_ROOT/Release/Models/CLIP-DataComp/metadata.json"
The SAM exporter downloads the gated checkpoint by model ID, exports a source
asset and an optimized runtime asset, and writes tokenizer and metadata. Select
the optimized sam3_float16.aimodel so the package contains the runtime asset
only. The selector refreshes the bundle fingerprint metadata.
cd "$PHOTOAIKIT_REPO"
export HF_HUB_OFFLINE=1
uv run Tools/export_sam3.py \
--model facebook/sam3 \
--dtype float16 \
--output-dir "$BUILD_ROOT/Release/Models" \
--bundle-name SAM3 \
2>&1 | tee "$BUILD_ROOT/Release/Evidence/sam3-export.log"
python3 Tools/select_sam3_asset.py sam3_float16.aimodel \
--bundle-dir "$BUILD_ROOT/Release/Models/SAM3"
Inspect the result. The package selector below includes only the optimized
asset, tokenizer and metadata; it does not include any remaining
*_source.aimodel directory.
test -f "$BUILD_ROOT/Release/Models/SAM3/metadata.json"
test -f "$BUILD_ROOT/Release/Models/SAM3/tokenizer/tokenizer.json"
test -f "$BUILD_ROOT/Release/Models/SAM3/sam3_float16.aimodel/main.mlirb"
cat "$BUILD_ROOT/Release/Models/SAM3/metadata.json"
7. Convert Qwen3-VL-2B-Instruct
Apple’s VLM exporter
creates the text decoder, embedding lookup, vision encoder, tokenizer, and
metadata.json in a qwen3_vl_2b directory. Use the full model: do not set
--num-layers or --skip-vision. The existing RawCull bundle records a 4096
token context. The exporter has no --revision option, so the isolated cache
and offline mode from section 4 matter here. Its output directory is the
parent of qwen3_vl_2b.
cd "$BUILD_ROOT/Tools/coreai-models"
export HF_HUB_OFFLINE=1
uv run coreai.vlm.export --list-models
uv run coreai.vlm.export qwen3-vl \
--max-context-length 4096 \
--compression none \
--output-dir "$BUILD_ROOT/Release/Models/Qwen" \
2>&1 | tee "$BUILD_ROOT/Release/Evidence/qwen-export.log"
If the pinned converter revision does not offer qwen3-vl or one of these
options, stop and inspect that revision’s --help; record and review any
converter revision change. Qwen export can take a long time and use substantial
memory. A process killed by macOS needs a fresh build directory or a carefully
inspected incomplete-output cleanup before retrying.
QWEN_BUNDLE="$BUILD_ROOT/Release/Models/Qwen/qwen3_vl_2b"
test -f "$QWEN_BUNDLE/metadata.json"
test -f "$QWEN_BUNDLE/tokenizer/tokenizer.json"
test -f "$QWEN_BUNDLE/qwen3_vl_2b.aimodel/main.mlirb"
test -f "$QWEN_BUNDLE/embed.aimodel/main.mlirb"
test -f "$QWEN_BUNDLE/vision.aimodel/main.mlirb"
python3 -m json.tool "$QWEN_BUNDLE/metadata.json"
Check that the metadata still names Qwen/Qwen3-VL-2B-Instruct, the three
assets, 448 px vision input, and 4096 context. This is a format check, not an
inference test. RawCull’s Qwen provider must also load and run the bundle.
8. Stage licences, notices and build provenance
Copy the reviewed notice catalog from RawCull. These files include the complete
model, tokenizer, and Apple conversion-recipe notices used by the existing
release. Review their terms and dates against the newly downloaded sources
before redistribution. SAM 3’s licence acceptance is required in RawCull.
for name in CLIP-DataComp SAM3 Qwen; do
ditto "$RAWCULL_REPO/ModelAssets/Notices/$name" \
"$BUILD_ROOT/Release/Notices/$name"
done
find "$BUILD_ROOT/Release/Notices" -type f -maxdepth 2 -print | sort
Those checked-in NOTICE.md files describe earlier published versions. Mark
the staging copies as new candidates without changing their licence text:
python3 - "$BUILD_ROOT" <<'PY'
from pathlib import Path
import sys
base = Path(sys.argv[1]) / 'Release/Notices'
replacements = {
'CLIP-DataComp': (
'RawCull publishes this pack in the v2 model release. The release catalog\n'
'records its download size and version, while the release host records the\n'
'archive checksum. This notice catalog records the upstream reference revision,\n'
'runtime fingerprint, and complete accompanying licence notices.',
'This converted bundle is a new local release candidate. Its final archive\n'
'size, SHA-256, Apple-assigned version, and review status must be recorded\n'
'after packaging and upload.'),
'SAM3': (
'The Apple-hosted asset pack is enabled for download at the project owner\'s\n'
'direction. Its archive byte size and SHA-256 are recorded in the external\n'
'release evidence after packaging, while the host-correct in-pack release record\n'
'is in `PROVENANCE.json`. This release decision does not claim an independent\n'
'legal review. Verified licence acceptance remains required.',
'This converted bundle is a new local release candidate. Record its archive\n'
'size, SHA-256, Apple-assigned version, and review status after packaging\n'
'and upload. Verified SAM licence acceptance remains required.'),
'Qwen': (
'RawCull publishes this pack in the v3 model release. The release catalog\n'
'records its download size and version, while the release host records the\n'
'archive checksum.',
'This converted bundle is a new local release candidate. Record its final\n'
'archive size, SHA-256, Apple-assigned version, and review status after\n'
'packaging and upload.'),
}
for name, (old, new) in replacements.items():
path = base / name / 'NOTICE.md'
content = path.read_text()
if content.count(old) != 1:
raise RuntimeError(f'Expected release paragraph not found: {path}')
path.write_text(content.replace(old, new))
PY
The copied PROVENANCE.json files describe the old release. Replace only
the staging copies with a clearly identified record of this build. The script
below preserves the licence inventory, writes the actual converted-component
hashes, and points to the source inventory. It intentionally has no final .aar
hash because that hash cannot be embedded inside its own archive.
python3 - "$BUILD_ROOT" "$CLIP_REV" "$SAM_REV" "$QWEN_REV" <<'PY'
import datetime, hashlib, json, pathlib, sys
root = pathlib.Path(sys.argv[1])
revisions = dict(zip(('CLIP-DataComp', 'SAM3', 'Qwen'), sys.argv[2:]))
models = {
'CLIP-DataComp': root / 'Release/Models/CLIP-DataComp',
'SAM3': root / 'Release/Models/SAM3',
'Qwen': root / 'Release/Models/Qwen/qwen3_vl_2b',
}
for name, model_dir in models.items():
notice_dir = root / 'Release/Notices' / name
old = json.loads((notice_dir / 'PROVENANCE.json').read_text())
components = {}
for path in sorted(model_dir.rglob('main.mlirb')):
digest = hashlib.sha256()
with path.open('rb') as stream:
for chunk in iter(lambda: stream.read(4 * 1024 * 1024), b''):
digest.update(chunk)
components[str(path.relative_to(model_dir))] = digest.hexdigest()
record = {
'catalog_version': 2,
'release_status': 'candidate',
'release': {
'hosting': 'apple',
'app_bundle_id': 'no.blogspot.RawCull',
'asset_pack_id': {
'CLIP-DataComp': 'rawcull-clip-datacomp',
'SAM3': 'rawcull-sam3',
'Qwen': 'rawcull-qwen3-vl-2b',
}[name],
'packaging_date': datetime.date.today().isoformat(),
'processing_status': 'not-uploaded',
'review_state': 'not-submitted',
},
'model': {
'bundle': old.get('model', {}).get('bundle', name),
'converted_main_mlirb_sha256': components,
},
'upstream': {
'project': old.get('upstream', {}).get('project'),
'selected_revision': revisions[name],
'source_inventory': 'Release/Evidence/source-files-sha256.txt',
'exporter_binding_note': 'Verify exporter log and isolated cache before claiming exact source binding.',
},
'conversion': {
'photoaikit_commit_file': 'Release/Evidence/photoaikit-commit.txt',
'coreai_models_commit_file': 'Release/Evidence/coreai-models-commit.txt',
},
'licences': old.get('licences', []),
}
(notice_dir / 'PROVENANCE.json').write_text(json.dumps(record, indent=2) + '\n')
PY
The staging provenance format is release-candidate evidence; it is not a
drop-in replacement for RawCull’s checked-in PROVENANCE.json. After packaging,
update the repository record using its full validated schema and the final
archive hash, size, App Store Connect pack version, and processing state.
9. Create the three packaging manifests
These selector paths are relative to the current directory used by
ba-package, which must be BUILD_ROOT/Release. Keep the permanent IDs and
installed paths identical to RawCull’s catalog.
cat > "$BUILD_ROOT/Release/Packaging/clip-datacomp.json" <<'JSON'
{
"assetPackID": "rawcull-clip-datacomp",
"downloadPolicy": { "onDemand": {} },
"fileSelectors": [
{ "file": "Models/CLIP-DataComp/metadata.json" },
{ "directory": "Models/CLIP-DataComp/tokenizer" },
{ "directory": "Models/CLIP-DataComp/ViT-B-32-256-datacomp_s34b_b86k_float16_static.aimodel" },
{ "directory": "Notices/CLIP-DataComp" }
],
"platforms": ["macOS"]
}
JSON
cat > "$BUILD_ROOT/Release/Packaging/sam3.json" <<'JSON'
{
"assetPackID": "rawcull-sam3",
"downloadPolicy": { "onDemand": {} },
"fileSelectors": [
{ "file": "Models/SAM3/metadata.json" },
{ "directory": "Models/SAM3/tokenizer" },
{ "directory": "Models/SAM3/sam3_float16.aimodel" },
{ "directory": "Notices/SAM3" }
],
"platforms": ["macOS"]
}
JSON
cat > "$BUILD_ROOT/Release/Packaging/qwen3-vl-2b.json" <<'JSON'
{
"assetPackID": "rawcull-qwen3-vl-2b",
"downloadPolicy": { "onDemand": {} },
"fileSelectors": [
{ "file": "Models/Qwen/qwen3_vl_2b/metadata.json" },
{ "directory": "Models/Qwen/qwen3_vl_2b/tokenizer" },
{ "directory": "Models/Qwen/qwen3_vl_2b/embed.aimodel" },
{ "directory": "Models/Qwen/qwen3_vl_2b/qwen3_vl_2b.aimodel" },
{ "directory": "Models/Qwen/qwen3_vl_2b/vision.aimodel" },
{ "directory": "Notices/Qwen" }
],
"platforms": ["macOS"]
}
JSON
for slug in clip-datacomp sam3 qwen3-vl-2b; do
python3 -m json.tool "$BUILD_ROOT/Release/Packaging/$slug.json" >/dev/null
done
.DS_Store files are present in the older ModelAssets/Release directories;
the new candidate should not include them. Do not delete anything from the old
release. Check the fresh selected directories and resolve any unexpected
symlink or secret before packaging.
cd "$BUILD_ROOT/Release"
find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
\( -name '.DS_Store' -o -type l \) -print
find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
-type f -print0 | xargs -0 shasum -a 256 | sort \
> Evidence/selected-inputs-sha256.txt
for slug in clip-datacomp sam3 qwen3-vl-2b; do
xcrun ba-package evaluate "Packaging/$slug.json" \
| tee "Evidence/$slug-evaluate.txt"
done
Read all three Evidence/*-evaluate.txt files. Each list should contain only
its model bundle, tokenizer, metadata, and matching notice directory. The
Qwen pack needs all three .aimodel directories. No source weight files,
download cache, old archive, or other model should be selected. If the
evaluation output is wrong, fix the manifest and rerun evaluation and the
input inventory before packaging.
11. Build the three .aar files
ba-package creates Background Assets archives; .aar is not a ZIP file.
Run it from BUILD_ROOT/Release so the relative file selectors resolve.
cd "$BUILD_ROOT/Release"
xcrun ba-package package Packaging/clip-datacomp.json \
--output-path Output/clip-datacomp.aar --verbose \
2>&1 | tee Evidence/clip-datacomp-package.log
xcrun ba-package package Packaging/sam3.json \
--output-path Output/sam3.aar --verbose \
2>&1 | tee Evidence/sam3-package.log
xcrun ba-package package Packaging/qwen3-vl-2b.json \
--output-path Output/qwen3-vl-2b.aar --verbose \
2>&1 | tee Evidence/qwen3-vl-2b-package.log
Do not edit an archive after this point. Any changed model, metadata, notice,
or manifest requires another ba-package package run and a new hash. For the
three existing App Store Connect pack records, a changed archive will become a
new pack version under the same permanent ID.
12. Verify and record the result
cd "$BUILD_ROOT/Release"
for slug in clip-datacomp sam3 qwen3-vl-2b; do
test -s "Output/$slug.aar"
stat -f '%N|%z bytes' "Output/$slug.aar"
shasum -a 256 "Output/$slug.aar"
shasum -a 256 "Packaging/$slug.json"
done | tee Evidence/archive-and-manifest-sha256.txt
find Output -maxdepth 1 -type f -name '*.aar' -print | sort
The last command must print exactly these three paths:
Output/clip-datacomp.aar
Output/qwen3-vl-2b.aar
Output/sam3.aar
Compare the final selected-inputs-sha256.txt with a fresh hash pass to catch
any source mutation while the archives were built:
find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
-type f -print0 | xargs -0 shasum -a 256 | sort \
> Evidence/selected-inputs-after-sha256.txt
diff -u Evidence/selected-inputs-sha256.txt \
Evidence/selected-inputs-after-sha256.txt
An empty diff and exit status 0 confirm that the selected inputs remained
unchanged during packaging. Record the BUILD_ROOT path, Xcode version,
exporter commits, source hashes, and all three archive hashes with the release
candidate. The files are at:
<BUILD_ROOT>/Release/Output/clip-datacomp.aar
<BUILD_ROOT>/Release/Output/sam3.aar
<BUILD_ROOT>/Release/Output/qwen3-vl-2b.aar
13. Before uploading or calling these release files
- Verify that the DataComp exporter actually consumed the pinned checkpoint;
the downloaded LAION snapshot alone does not prove it. For all three packs,
retain actual source-weight hashes and exporter logs.
- Review the current model licences and all copied notice files. Keep the
correct notice directory inside each
.aar. - Run RawCull’s
make verify-model-provenance, catalog and release-metadata
tests, and release preflight after updating its manifest template,
Swift catalog, and checked-in provenance to the new archive values. Never
leave the old archive SHA-256 in the app catalog for new .aar files. - If uploading, use the existing
rawcull-clip-datacomp, rawcull-sam3, and
rawcull-qwen3-vl-2b App Store Connect records. Follow
RawCull/Docs/newmodels.md
for xcrun altool, API-key handling, processing checks, and TestFlight
verification. The AppStore build must use Apple hosting and the matching
App Group. A successful local .aar build does not test runtime inference.
If a step fails
| Symptom | Check |
|---|
401/403 downloading SAM 3 | Hugging Face approval and hf auth whoami; accept the gated model terms. |
| Offline source missing | Check the isolated HF_HUB_CACHE and the model’s refs/main; download the exact revision before re-exporting. |
| CLIP tokenizer parity failure | Do not package; inspect the tokenizer source and OpenCLIP configuration. |
| SAM export leaves source asset | Package only the optimized asset selected by select_sam3_asset.py. |
Qwen bundle lacks vision.aimodel | Rebuild with a converter that supports Qwen VLM and without --skip-vision. |
ba-package evaluate lists extra files | Correct its fileSelectors; evaluate again before packaging. |
| Archive hash differs from the old release | Expected for a new conversion; update catalog and provenance before upload. |
The source commands and pack layout come from PhotoAIKit’s export tools,
Apple’s Core AI model recipes,
Apple’s managed pack documentation,
and the RawCull release runbook linked above. Verify the exact checkout used
for a release: repository main branches and tool versions can change.
10.2 - AI Model Licence and Provenance Clearance
Current catalog status and historical model-clearance evidence.
AI model licence and provenance clearance procedure
Current code status (September 26, 2026): the 3.2.6 production catalog
enables Apple-hosted DataComp CLIP, Meta SAM 3, and Qwen3-VL-2B-Instruct.
The Direct/Developer ID configuration retains the historical self-hosted v3
manifest. OpenAI CLIP is excluded and EfficientSAM is not a production pack.
The detailed clearance evidence below is a September 15 historical snapshot
for the earlier two-pack self-hosted release; it has not been re-reviewed as a
legal or provenance opinion on the current Apple-hosted Qwen pack. For current
pack IDs, hashes, licences, and test-release steps, see
Publishing and Testing RawCull AI Models
and the application repository ModelAssets/README.md.
Technical repository evidence in the historical snapshot: 2026-09-15
Evidence record owner: Thomas Evensen, RawCull maintainer
September 26 code snapshot
| Production pack | Recorded licence | Explicit in-app acceptance | Current evidence location |
|---|
| DataComp CLIP | OpenCLIP/DataComp MIT notice | No | Catalog descriptor and ModelAssets/Notices/CLIP-DataComp |
| Meta SAM 3 | SAM License, November 19, 2025 | Yes, with a verified bundled text | Catalog descriptor and ModelAssets/Notices/SAM3 |
| Qwen3-VL-2B-Instruct | Apache License 2.0 | No | Catalog descriptor and ModelAssets/Notices/Qwen |
The three descriptors are .ready in the current code and their archive
checksums are recorded in the model release guide. This table
reports repository metadata, not independent legal clearance or an App Review
decision. Reassess the notices and provenance for the exact packs submitted.
Historical distribution-status snapshot (September 15, 2026)
This section is a dated status snapshot. It describes the September 15 product
and repository records; it is not a legal conclusion and must not be copied into
a later release without a fresh evidence review.
| Pack | Current product/release record | Evidence record | Residual point for the next publication | Owner/action before next publication |
|---|
| DataComp CLIP | .ready, enabled, published in v3 | Archive size/SHA-256, runtime fingerprint, reference revision, tokenizer and notice hashes are recorded | Provenance still records the upstream revision as a reference and leaves source_weight_sha256 null | Bind the exact weight file on a rebuild or preserve a signed residual-provenance decision |
| OpenAI CLIP | .ready in the prepared catalog, excluded from production | Historical v2 archive and pinned source evidence remain recorded | The weight-specific licence basis described below remains a future-publication question | Reassess and record a named approval before enabling it again |
| Meta SAM 3 | .ready, enabled, published in v3; verified licence acceptance required | Archive size/SHA-256, source revision/checksum, runtime hash, complete licence and notice hashes are recorded | The upstream checkpoint is gated; the repository records the project owner’s release decision, not an independent legal opinion | Preserve the decision and evidence; reopen review if terms, delivery, model, or licence text changes |
| EfficientSAM | .blocked in the prepared catalog, excluded from production | Source/checkpoint/conversion/licence metadata are prepared | Final converted fingerprint and archive size/SHA-256 are absent | Keep excluded until its descriptor and provenance pass the complete gate |
The application catalog and ModelAssets records are the authoritative account
of what that self-hosted v3 release contained: DataComp and SAM 3.
OpenAI CLIP and EfficientSAM do not pass the inclusion flags into the production
catalog or manifest template. A .ready value proves only
that the product gate was opened. Model availability, a public archive, or a
model-page licence badge must never be treated alone as permission for the
specific conversion and redistribution.
Purpose of the reusable procedure
This document defines how RawCull clears the DataComp CLIP, OpenAI CLIP, and
Meta SAM 3 model packs for public download. It covers technical provenance,
licence evidence, upstream contacts, questions to ask, acceptable answers, and
the final release gate.
For a new publication, do not upload a new archive, publish a new download
manifest, or change a production descriptor to ready until every candidate is
either:
- cleared under this procedure; or
- deliberately excluded from the release and manifest.
Existing release archives are immutable historical evidence. The safest
remedy for a recorded provenance gap is a new candidate and new release tag
created from pinned, hashed source files after the applicable licence decision
has been reviewed; do not rewrite history by silently replacing the old record.
This is an engineering and evidence-preservation procedure, not legal advice.
For an unresolved interpretation, obtain advice from a qualified lawyer who
works with software copyright, open-source licensing, AI model weights, and
commercial distribution in Norway and the EEA.
Reusable clearance standard
What “resolved” means
There are two independent gates.
Provenance gate
RawCull must be able to demonstrate this complete chain:
upstream owner and repository
↓
immutable revision and exact source-weight filename
↓
SHA-256 of the downloaded source file
↓
pinned conversion code, command, dependencies, and toolchain
↓
SHA-256 of the converted runtime model
↓
SHA-256 and byte size of the packaged Managed Background Assets pack
A filename, a local cache timestamp, or a likely upstream snapshot is not a
cryptographic binding. Re-exporting from a deliberately selected and hashed
source is preferable to trying to infer the origin of an old conversion.
Licence and distribution gate
RawCull must have a defensible basis for all of the following:
- the licence applies to the exact trained weight file, not only source code;
- conversion into an Apple Core AI model is allowed;
- the converted derivative can be redistributed to third parties;
- the intended distribution can be public and commercial, if RawCull is
commercially distributed;
- all notices, agreement copies, acceptance steps, use restrictions, and
attribution requirements have been implemented; and
- a gated source checkpoint may be redistributed through RawCull’s proposed
delivery mechanism, if applicable.
Silence, an unanswered ticket, a community member’s assumption, widespread
third-party mirroring, or the technical ability to download a file does not
resolve this gate. Prefer a written response from the model owner or an
authorized representative. If that is unavailable, obtain a written opinion
from qualified counsel or omit the model.
Current recorded evidence
The canonical records are:
ModelAssets/Notices/CLIP-DataComp/PROVENANCE.jsonModelAssets/Notices/CLIP-OpenAI/PROVENANCE.jsonModelAssets/Notices/SAM3/PROVENANCE.json- the licence and notice files beside each provenance record
The current relevant identifiers are:
| Pack | Upstream revision presently recorded | Source-weight evidence | Converted runtime evidence | Open issue |
|---|
| DataComp CLIP | 4afec35ffe57a943d569ff7ee888061830164da8 is a reference revision, not exporter-recorded proof | Exact selected source-weight SHA-256 remains null in provenance | runtime main.mlirb SHA-256 41596f6f7a9f8f8d1171b0056f4e3a90902ef88d73303713ab3bed4847b6266d; directory fingerprint 6a3639a2049b8a4ea23fe04c3083e199a4f505433f7c8bd0748b3c8d4fcb1572 | Bind the exact weight input on the next rebuild and recheck licence/model-card terms |
| OpenAI CLIP | 3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268, also recorded by the exporter | pytorch_model.bin, SHA-256 a63082132ba4f97a80bea76823f544493bffa8082296d62d71581a4feff1576f | runtime main.mlirb SHA-256 828e6ef52700c48b9c72696d785f5b45bd01a06748c538eb284cf3a42f2530da; directory fingerprint 24a20d7c5c88da2afe3ed81dca0ddf223450dd6afd1f3aff34be7acfc48f4914 | Preserve the named approval basis for weight redistribution; re-open the gate if that evidence is missing or changes |
| SAM 3 | 3c879f39826c281e95690f02c7821c4de09afae7; not exporter-bound | model.safetensors, SHA-256 6d06f0a5f84e435071fe6603e61d0b4cc7b40e0d39d487cfd4d67d8cc11cc14a | runtime main.mlirb SHA-256 43a9b88e40d193f5a6608a7fee536a78f4ba4ec5d95f1eb24db03031630f0a31 | Confirm whether an ungated public derivative download is compatible with Meta’s gated access flow, then rebuild with the current exporter |
These hashes identify current evidence; they do not themselves approve a
release.
Perform these steps separately for every model that passes its licence gate.
- Choose one exact upstream repository, immutable commit, and source-weight
file. Never use
main, latest, an unpinned model alias, or an automatically
changing download URL as the release input. - Download into a new, dated evidence directory. Preserve the upstream URL,
immutable revision, filename, byte size, and SHA-256 before conversion.
- Save the model card, licence or agreement, repository metadata, and any
access terms as they appeared on the download date. Record their URLs and
retrieval dates.
- Record the converter repository and commit, Apple
coreai-models commit,
coreai-core version, Python environment or package lock, conversion
command, macOS version, Xcode version, and conversion timestamp. - Run the conversion from that evidence directory. Do not allow the exporter
to resolve or download a floating model identifier internally.
- Hash the complete converted model directory with the established
directory-tree-sha256-v1 method and hash its runtime main.mlirb file. - Validate the converted model with the same PhotoAIKit checks used by
RawCull.
- Update the corresponding
PROVENANCE.json with the actual source revision,
source filename, source SHA-256, conversion command or record, tool versions,
output fingerprints, and supporting evidence references. - Rebuild the extensionless asset pack using the explicit selectors. Verify that the chosen runtime
model, tokenizer,
metadata.json, and complete notice catalog are present,
and that _source.aimodel and conversion intermediates are absent. - Record the new asset-pack byte size and SHA-256. The previous unpublished archive hash
must not be reused for the rebuilt archive.
An example pinned Hugging Face acquisition has this form; select the correct
source filename before running it:
hf download OWNER/MODEL SOURCE_WEIGHT_FILE \
--revision IMMUTABLE_COMMIT \
--local-dir /path/to/private/release-evidence/MODEL/source
shasum -a 256 \
/path/to/private/release-evidence/MODEL/source/SOURCE_WEIGHT_FILE
Keep raw correspondence, access tokens, account data, and legal advice out of
the public repository. Store them in private, access-controlled records. A
public provenance summary may record the response date, organization, scope,
and internal evidence-record identifier without publishing personal data or
privileged legal advice.
DataComp CLIP clearance
Current position
RawCull currently records this pack as ready and publishes it in v3. The
pinned DataComp repository page reviewed on 2026-08-22 identifies the checkpoint
as MIT licensed, which is positive evidence. The remaining technical gap is
that the packaged provenance does not identify and hash the exact source-weight
file used by the exporter. That gap must remain visible in the decision record;
the published state does not make it disappear.
Official references:
- Email LAION at
contact@laion.ai. LAION publishes this address on its
official legal contact page. - Open a discussion on the exact Hugging Face model repository so the question
and any maintainer response are tied to that checkpoint.
- If necessary, open an issue with the OpenCLIP maintainers to identify which
upstream file the
datacomp_s34b_b86k configuration resolves. OpenCLIP can
help with technical identity; the model owner or counsel should resolve
licence scope.
Recorded LAION email request
On 2026-08-02 at 17:20 CEST, the RawCull maintainer emailed
contact@laion.ai with the subject “Written clarification requested for
DataComp CLIP weight licensing and redistribution.” The request identified:
- repository:
laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K; - immutable revision:
4afec35ffe57a943d569ff7ee888061830164da8; - proposed source-weight file:
open_clip_model.safetensors; - model configuration: OpenCLIP
ViT-B-32-256 with
datacomp_s34b_b86k weights; - transformation: a float16 Apple Core AI runtime representation for local
photo similarity and text-to-image semantic search; and
- delivery: an optional model archive downloaded by RawCull from a public
GitHub Release, including possible use with a commercially distributed
version of RawCull.
The email stated that RawCull has not publicly distributed the model or its
converted derivative and will not redistribute the DataComp training dataset.
It proposed preserving the source byte size and SHA-256 and including the
applicable MIT licence, copyright and attribution notices, model information,
provenance, and checksums with the archive.
LAION was asked to confirm whether the displayed MIT licence covers the exact
weight file, conversion into the proposed runtime representation, public and
commercial redistribution of the derivative, and whether any additional
DataComp, OpenCLIP, attribution, acceptable-use, or other conditions apply. The
request also asked whether the model card’s “out of scope” deployment language
is safety guidance or an additional legal restriction.
Status: awaiting a substantive response from LAION or another person
authorized to clarify the rights applicable to the weights. Sending the email
does not clear the pack. Preserve the original message and any response in the
private evidence register; do not add the maintainer’s personal email address
or complete mail headers to this public repository.
Questions to ask
Identify the exact repository, revision, and proposed source file, then ask:
- Is that exact trained weight file offered under the MIT licence shown on the
model repository?
- Does that permission cover conversion into another runtime representation
and redistribution of the converted weights with a desktop application?
- Is public and commercial redistribution permitted, provided the MIT notice
is included?
- Are any DataComp dataset terms, OpenCLIP terms, attribution requirements, or
use restrictions additional to the displayed MIT licence applicable to the
weight file?
Required technical work
- Select exactly one of the upstream weight files rather than leaving the
exporter to resolve a model alias.
- Download it at revision
4afec35ffe57a943d569ff7ee888061830164da8 and record its SHA-256 and byte
size. - Re-export DataComp CLIP from that local file under the common procedure.
- Replace the null source revision and source checksum in
ModelAssets/Notices/CLIP-DataComp/PROVENANCE.json.
Sufficient resolution
The pack may pass this gate when both conditions hold:
- the new conversion has a complete pinned-and-hashed provenance chain; and
- LAION or another demonstrably authorized model owner confirms the licence
scope in writing, or qualified counsel concludes in writing that the
repository’s MIT designation and accompanying materials are sufficient for
the intended distribution.
If the answer is negative or remains materially ambiguous, replace the model
with a checkpoint having explicit weight-level redistribution terms or omit
the DataComp pack.
DataComp CLIP - no answer
If LAION does not provide a substantive answer after a documented follow-up,
silence neither grants additional permission nor withdraws the permission
already stated in the published materials. A release under the displayed MIT
licence would rely on the public licence evidence rather than individualized
clearance from LAION. Record it as a maintainer risk-acceptance decision, not as
an upstream-approved or upstream-cleared release.
The evidence supporting that decision is:
- the exact pinned model repository identifies the model as
License: mit and
contains the weight files; - Hugging Face’s licence documentation
describes model-card licence metadata as communicating the permissions
attributed to repository content; and
- the MIT licence permits use,
modification, publication, redistribution, sublicensing, and sale when its
copyright and permission notice accompanies copies or substantial portions.
The residual ambiguity is that the model repository has MIT metadata but no
standalone LICENSE file identifying the trained weights and their copyright
holder. The OpenCLIP MIT notice clearly covers the OpenCLIP software, but an
unanswered inquiry leaves no individualized confirmation that it is also the
intended notice for the trained weights and converted derivative. The model
card’s “out of scope” deployment language appears as safety and intended-use
guidance rather than licence text, but this interpretation has not been
confirmed by LAION. Qualified Norwegian counsel remains the recommended way to
resolve that ambiguity before a public or commercial release.
If the maintainer nevertheless decides to release without an answer, complete
all of the following before publication:
- Send and preserve one documented follow-up to LAION. Record the original
request, follow-up date, response deadline, and absence of a substantive
answer in the private evidence register.
- Select
open_clip_model.safetensors from revision
4afec35ffe57a943d569ff7ee888061830164da8 as the only conversion input. Its
pinned Hugging Face metadata reports a byte size of 605189364 and SHA-256
92c26d60d3200ed5ed040dff31a8d19f8140648da8007216c25744c478deef27.
Download the file independently and verify both values before conversion. - Re-export the Apple Core AI model from that pinned local file. Do not merely
add the upstream hash to the provenance record for the existing conversion;
the released derivative must be cryptographically tied to the verified
source file.
- Preserve dated copies of the pinned repository tree, model card, repository
API metadata, MIT licence, OpenCLIP notice, conversion command, dependency
versions, and all input and output checksums.
- Package the complete applicable MIT copyright and permission notice,
OpenCLIP and tokenizer notices, model card, safety limitations, provenance,
and conversion information with every redistributed model archive. A link
alone is not a substitute for including the required notice.
- Keep DataComp CLIP identified as a separate third-party model asset.
RawCull’s own MIT licence does not relicense the model, and RawCull must not
claim ownership of or permission to redistribute the DataComp training
dataset.
- Complete the common provenance procedure, PhotoAIKit validation, asset-pack
inspection, archive hashing, manifest verification, and download tests.
Change
PROVENANCE.json and the production catalogue to ready only after
those technical controls describe the new release candidate accurately. - Add a signed and dated release decision stating the evidence relied upon,
the unresolved licence ambiguity, whether RawCull is free or commercial,
the intended distribution countries and channels, and the responsible
maintainer’s acceptance of the residual risk.
- Keep the model pack independently removable from the manifest and release
so distribution can be suspended promptly if a credible rights claim or
contrary clarification is received.
Releasing after these steps may provide a defensible MIT-compliance position,
but it does not eliminate the residual legal risk created by the absence of a
weight-specific licence file or an authorized response. This procedure records
the evidence and decision; it is not a legal opinion.
OpenAI CLIP clearance
Current position
RawCull retains this pack as ready in the prepared catalog, but excludes it from
the production catalog and v3 manifest. The historical v2 record remains
relevant if this model is considered for a future release. The
OpenAI CLIP source repository contains an MIT licence covering the software and
associated documentation. The Hugging Face checkpoint page reviewed on
2026-08-22 still does not display a clear weight-level licence designation. A
community assumption that the repository licence covers the weights is not
authoritative evidence.
The model card also characterizes deployed uses as out of scope. Clarify
whether this is safety guidance or an enforceable distribution/use condition,
and independently assess RawCull’s use, testing, limitations, and disclosures.
Official references:
- Contact OpenAI Support using the chat control at
help.openai.com. Ask that
the request be routed to the team responsible for CLIP/open-source model
licensing. Retain the ticket or conversation identifier. - Submit the same narrowly framed question through the feedback form linked by
the official CLIP model card.
- Open a model-specific discussion on the
Hugging Face Community page
using its
New discussion
action. This requires a Hugging Face login. Use a discussion rather than a
pull request so the licensing question remains attached to the exact model
distribution.
- A response from an OpenAI employee or repository maintainer authorized to
address the model’s licence is preferred. An unsupported answer from another
community member is not sufficient.
GitHub issue creation for openai/CLIP is currently restricted, so it should
not be the only planned contact route.
Recorded support outcome and Hugging Face escalation
OpenAI Support declined to confirm or provide an authoritative interpretation
that the MIT licence in openai/CLIP applies to the pretrained weight files in
openai/clip-vit-base-patch32. Support also declined to confirm that the
licence permits RawCull’s intended commercial redistribution. The response
said that the applicable terms must be determined from the licence and notices
shipped with the exact code and weights being used.
This historical response did not by itself clear the pack. RawCull’s later
product record nevertheless marks the pack ready, so the private decision
register must identify the subsequent evidence, responsible approver, date, and
accepted scope. If it cannot, treat that as a blocker for the next publication
rather than pretending the historical uncertainty was resolved. The conclusion
at the time of the support exchange was:
CLIP source code: MIT licensed.
openai/clip-vit-base-patch32 pretrained weights: no explicit weight-specific
licence identified in the downloaded distribution, and OpenAI Support did not
confirm that the source-repository MIT licence applies.
Commercial redistribution: unresolved and blocked pending sufficient evidence.
The OpenAI Report Content form is intended for reports of potentially illegal
or policy-violating content. It is not evidence of permission and should not be
treated as the route for prospective licensing clearance. Support’s recommended
next step was to ask the maintainers on the distribution source. For Hugging
Face, use the model-specific Community page linked above rather than the general
Hugging Face forum.
On 2026-08-02, the RawCull maintainer opened
Hugging Face discussion #72,
“License applicable to pretrained CLIP ViT-B/32 weights,” asking the OpenAI
maintainers to identify the licence covering the hosted weights and to confirm
whether it permits commercial use and redistribution. The request also asks
the maintainers to add the applicable licence identifier or licence file to the
model repository. Opening the discussion records the escalation but does not
clear the pack; retain the response and assess the responder’s authority and
the scope of any answer before changing the release decision.
Suggested discussion title:
Licence applicable to pretrained CLIP ViT-B/32 weights
Suggested discussion body:
Could the OpenAI maintainers clarify the licence applicable specifically to
the pretrained weight files in openai/clip-vit-base-patch32?
In particular, does OpenAI intend the MIT License from the official
openai/CLIP repository to cover the original ViT-B/32 checkpoint and the
converted weight files hosted here, including commercial use, conversion to
another runtime representation, and redistribution subject to the MIT
conditions?
If so, could you add the applicable licence identifier and/or LICENSE file to
this model repository so downstream users can establish a reliable licensing
record?
A maintainer comment may clarify intent, but the strongest resolution is a
licence file, model-card licence declaration, or written statement from the
rights holder that expressly covers the exact weights and proposed use. Until a
responsible approver records the basis that supersedes this historical finding,
do not use existing v2 availability as the sole basis for a new commercial
redistribution decision.
Questions to ask
Provide this exact identity:
- repository:
openai/clip-vit-base-patch32; - revision:
3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268; - file:
pytorch_model.bin; - SHA-256:
a63082132ba4f97a80bea76823f544493bffa8082296d62d71581a4feff1576f.
Ask:
- What licence governs this exact trained weight file?
- Does the OpenAI CLIP MIT licence apply to it, or is there a separate licence
or set of terms?
- May RawCull convert the weights into an Apple Core AI representation and
publicly redistribute that converted derivative as an optional desktop-app
download, including in a commercial application?
- Which copyright notice, attribution, model card, use limitation, or other
terms must accompany the derivative?
- Is the model card’s statement that deployed uses are out of scope guidance,
or does OpenAI intend it as a legal restriction on deployment or
redistribution?
Required technical work
After licence clearance, re-export from the exact local source file and
revision above. Record the conversion command and bind the new output to the
source SHA-256. Update the provenance record and rebuild the extensionless asset pack.
Sufficient resolution
The pack may pass this gate only with:
- an authoritative written statement identifying terms that permit the exact
intended conversion and redistribution; or
- a written legal opinion accepting a clearly documented alternative basis.
A response that merely restates the code repository’s MIT licence without
addressing the trained weights is not sufficient. If OpenAI does not answer or
the answer does not permit the proposed distribution, omit OpenAI CLIP or use a
replacement checkpoint with explicit weight-level terms.
Current position
The SAM License dated November 19, 2025 defines the trained weights as SAM
Materials and grants rights to use, reproduce, distribute, modify, and create
derivatives. Distribution must remain under that agreement and a copy of the
agreement must accompany the materials. RawCull now packages the complete
agreement and requires acceptance of its verified text.
The separate unresolved question arises because Meta’s official checkpoint is
gated. The model page requires a user to log in and share contact information,
and Meta’s repository instructs users to request access and authenticate before
downloading. Hugging Face documents that a gated model’s authors control
access. The licence text appears permissive about redistribution, but a public
GitHub download would let downstream users obtain the derivative without going
through Meta’s upstream access flow.
The repository’s current product decision nevertheless marks SAM 3 ready,
enables it in production v3, and records verified archive metadata. This
technical state does not erase the review topic below or claim independent
legal advice; preserve the decision record and reopen it when the model,
licence, delivery mechanism, or upstream terms change.
Official references:
- Email Meta Open Source at
opensource@meta.com. Meta publishes this contact
address on its official Open Source terms page. Ask that the request be
routed to the SAM 3 model/licensing owner. - Open a narrowly scoped issue in
facebookresearch/sam3 and/or a discussion
on facebook/sam3. Link the public question from the private email so the
project team can answer in its preferred channel. - Contact Hugging Face only for clarification of how its gate operates.
Hugging Face is the hosting platform; its support cannot substitute for
permission or interpretation from Meta as the model owner.
- If Meta does not give a clear response, obtain a written opinion from
qualified counsel or omit SAM 3 from public hosting.
Provide this exact identity and delivery proposal:
- repository:
facebook/sam3; - revision:
3c879f39826c281e95690f02c7821c4de09afae7; - file:
model.safetensors; - SHA-256:
6d06f0a5f84e435071fe6603e61d0b4cc7b40e0d39d487cfd4d67d8cc11cc14a; - transformation: float16 Apple Core AI derivative for local inference;
- delivery: optional RawCull Managed Background Asset from a public GitHub
Release;
- licence handling: complete SAM License packaged with the derivative and
explicit in-app acceptance tied to the licence-text SHA-256.
Ask:
- Does section 1 of the SAM License permit this converted derivative to be
distributed from a public, ungated GitHub Release?
- Must every downstream RawCull user independently request access through
Meta’s Hugging Face gate before receiving the converted derivative?
- If downstream gating is required, what information and approval must
RawCull collect, and may RawCull technically administer that gate?
- Is packaging the complete agreement and requiring verified in-app
acceptance sufficient to distribute under the same agreement?
- Are there additional attribution, branding, reporting, geographic, trade
control, or prohibited-use measures RawCull must implement?
- Does the answer apply to both free and commercial distribution of RawCull?
Possible outcomes
Meta confirms public ungated redistribution. Preserve the response, comply
with every stated condition, re-export from the pinned source, update the
provenance record, and have counsel review any material ambiguity before
release.
Meta requires each recipient to pass its gate. Do not publish SAM 3 as a
public GitHub Release asset. Managed Background Assets with an anonymous public
origin will not reproduce individualized Hugging Face approval. Omit SAM 3 or
design a separate authenticated acquisition flow and review it independently.
Meta declines, gives an unclear answer, or does not respond. Keep SAM 3
blocked. Obtain a legal opinion or omit it. Do not treat the licence’s general
redistribution wording as resolving a separately imposed access condition
without an accountable decision.
Required technical work after clearance
Re-export from the recorded revision and model.safetensors SHA-256, recording
the complete conversion command and environment. Preserve the full SAM License
with the pack, maintain explicit acceptance against its checksum, and rebuild
the extensionless asset pack. Re-check the official licence immediately before publication because
the agreement states that Meta may modify it.
Use a real name and reply-capable address. Do not send an archive, proprietary
source, access token, or confidential material unless the recipient requests it
through an appropriate channel.
Subject: Written clarification requested for redistribution of [MODEL] in RawCull
Hello,
I maintain RawCull, a macOS photo-culling application:
https://github.com/rsyncOSX/RawCull
I have not published or uploaded the model derivative described below. I am
seeking written clarification before doing so.
Upstream repository: [URL]
Immutable revision: [COMMIT]
Source weight file: [FILENAME]
Source SHA-256: [SHA256]
RawCull converts this file to a float16 Apple Core AI model for local inference.
The proposed delivery is an optional Managed Background Asset downloaded from a
public GitHub Release. The archive will include the applicable complete licence,
notices, attribution, provenance, and checksums. [Describe explicit acceptance,
if applicable.] RawCull is [free/paid; choose the accurate description].
Please confirm:
1. Which licence or terms govern this exact trained weight file?
2. May it be converted into this runtime representation?
3. May the converted derivative be publicly redistributed in this manner,
including as part of a commercial application if applicable?
4. What notices, acceptance flow, attribution, gating, use restrictions, or
other conditions must RawCull implement?
I would appreciate a response from the model owner or a person authorized to
clarify its licensing and distribution conditions. If another team handles
this request, please route it or identify the correct contact.
Thank you,
[NAME]
[CONTACT DETAILS]
For SAM 3, append the six SAM-specific questions above. For OpenAI CLIP,
append the exact question about whether the MIT licence covers the trained
weights and whether the model card’s deployment language is guidance or a
restriction. For DataComp, ask whether the repository’s MIT designation covers
the exact selected file.
When to involve a lawyer
Engage a Norwegian lawyer before publication if any of these apply:
- the model owner does not answer;
- the answer is informal, conditional, or ambiguous;
- a licence permits redistribution but an upstream gate suggests a different
access policy;
- commercial use, downstream acceptance, trade controls, prohibited uses, or
indemnification requires interpretation;
- the responder’s authority to bind the rights holder is uncertain; or
- RawCull will distribute in multiple jurisdictions.
Look for counsel with experience in copyright, software and open-source
licensing, technology transactions, AI model weights, EEA law, and US/EU trade
controls. The
Norwegian Bar Association member search
can verify professional membership. The Norwegian Industrial Property Office
also publishes guidance on
choosing an IP adviser.
Give counsel a private evidence bundle containing:
- the exact model card and licence captured on the download date;
- upstream repository, revision, source filename, SHA-256, and download method;
- the conversion recipe and explanation of what the derivative contains;
- complete notice catalogs;
- every upstream response with message headers, dates, ticket IDs, and links;
- RawCull’s intended countries, free or paid status, distribution channels,
end-user flow, and licence-acceptance UI; and
- the proposed GitHub release and Managed Background Assets architecture.
Ask for a written conclusion for each model covering conversion,
redistribution, commercial use, notices, downstream terms, gating, and any
required technical controls. Keep privileged advice private. Record only the
resulting release decision and non-confidential obligations in the repository,
unless counsel approves broader disclosure.
Evidence and decision record
Maintain a private register with at least these fields:
| Field | Required content |
|---|
| Model | Exact upstream owner and repository |
| Source identity | Immutable revision, filename, byte size, SHA-256 |
| Licence identity | Name/version, official URL, captured file SHA-256, retrieval date |
| Contact record | Organization, channel, date, ticket/issue ID, responder and stated authority |
| Permission scope | Conversion, derivative redistribution, public access, commercial use, territories |
| Conditions | Notices, attribution, acceptance, gating, use restrictions, trade controls |
| Legal review | Counsel, date, private matter/reference number, approved/blocked conclusion |
| Conversion | Script/commit, command, dependencies, toolchain, timestamp, output hashes |
| Pack | Explicit selector manifest, asset-pack byte size and SHA-256, notice verification |
| Decision | Ready, blocked, replaced, or omitted; responsible approver and date |
Do not mark a contact item complete merely because a message was sent. Record
the actual answer and whether it addresses the exact file and proposed delivery.
Final release checklist
A pack can change from blocked to ready only when every applicable item is
complete:
If even one required item remains open, keep that pack blocked or omit it.
Publication sequence after clearance
RawCull may publish a ready subset. Every candidate must still have an explicit
ready, blocked, replaced, or omitted decision, and every blocked/omitted asset
must be absent from that manifest. The current v3 record follows this rule by
publishing DataComp CLIP and SAM 3 while excluding OpenAI CLIP and EfficientSAM.
After the decisions are complete:
- Re-export every approved model from its pinned and hashed source.
- Update the notice catalogs and change only genuinely approved catalogue
descriptors to
ready. - Rebuild and inspect the extensionless asset packs; record their new hashes and sizes.
- Generate and inspect the self-hosted download manifest with a non-beta or
corrected
ba-package toolchain. - Create the dedicated
RawCull-AI-Models release as a draft and upload only
approved packs, their required remote asset names, the manifest, and public
evidence. - Verify every URL, redirect, byte size, and checksum while authenticated to
the draft if necessary.
- Run download, acceptance, validation, removal, and licence-change tests
against a non-production environment.
- Publish the release only after a final human review confirms that the
manifest contains no blocked model and all obligations are satisfied.
Official-source summary
The following official sources were rechecked on 2026-08-22:
At that review date the DataComp page displayed License: mit; the OpenAI
checkpoint page did not display a corresponding weight-specific licence label;
and the official SAM 3 checkpoint remained gated with the SAM License linked.
These observations are evidence inputs, not legal conclusions.
Recheck all licence text, model-card metadata, gates, contacts, and official
links immediately before release because they can change.
10.3 - AI Models in RawCull
Code-level guide to local CLIP, Vision, SAM 3, Qwen, Deep Review, and numbered Objects analysis in RawCull.
AI Models in RawCull
RawCull uses several local machine-learning backends, but it does not treat them
as interchangeable. Each model family has a deliberately narrow job:
| Model or backend | RawCull job | Output used by RawCull |
|---|
| DataComp CLIP | Image similarity, burst grouping, semantic search, and coarse subject labels for Deep Review | Normalized image/text embedding vectors and cosine distances/similarities |
| OpenAI CLIP | Fully implemented alternative CLIP bundle; currently excluded from the production model list | The same typed CLIP artifacts as DataComp, with a different model fingerprint |
| SAM 3 | Prompted subject segmentation for Deep Review and separate instance segmentation for Objects | A chosen subject mask, or up to eight numbered masks per concept |
| Qwen3-VL-2B-Instruct | Standalone photo assessment; concept discovery and board interpretation in Objects | A photo assessment, validated object concepts and per-object findings, or a visible retryable response failure |
| Apple Vision feature print | Always-available image-similarity fallback | Opaque Vision feature-print artifacts and native distances |
All inference stays in the application process. The downloadable model assets
are installed separately because they are large, but the analysis path does not
send photographs to a remote inference service.
This page was checked against the local RawCull version-3.2.6 working tree
on September 26, 2026, including the packaged-model evaluation and later
in-app Objects observations. RawCull source links follow that branch.
PhotoAIKit links use revision
77cc1d84,
pinned by this checkout. The local working tree may be ahead of the published
branch until those changes are pushed.
Source Catalog: RawCull/Intelligence
The Intelligence directory is organized by responsibility rather than by one
model per folder. Start with this catalog when tracing the code.
Composition and contracts
Model management
| Source | Responsibility |
|---|
RawCullAIModelDownloadCatalog.swift | Production model inventory, inclusion switches, asset-pack identifiers, versions, byte counts, checksums, licences, and provenance links. |
RawCullAIModelDownloadService.swift | Background Assets download coordination and installed-location resolution. |
RawCullAIModelDownloadsModel.swift | Observable download, licence, progress, removal, and installed-location state. |
RawCullAIModelResourceManager.swift | Actor-isolated validation and provider construction with a metadata snapshot cache. |
RawCullAISettingsModel.swift | Applies installed locations, refreshes capabilities, stores user selections, and publishes revisioned runtime configurations. |
CLIP, similarity, and semantic search
| Source | Responsibility |
|---|
RawCullVisionSimilarityService.swift | Defines the shared similarity-service boundary, Vision implementation, CLIP implementation, RAW decoding adapter, finite-vector recovery, and artifact validation. |
SimilarityScoringModel.swift | Owns indexed artifacts, hydration, persistence, image ranking, grouping, semantic-search state, and CLIP-based subject classification. |
RawCullSimilarityFeature.swift | Stable application-facing similarity surface with cancellation and generation gates. |
RawCullSemanticSearchService.swift | Encodes a text query, admits compatible CLIP artifacts, compares image and text vectors, and ranks deterministically. |
RawCullSemanticSearchFeature.swift | Presentation and application-target adapter for semantic search. |
Deep Review with SAM 3 and CLIP
Objects: Qwen discovery, SAM 3 instances, Qwen review
| Source | Responsibility |
|---|
ObjectAnalysis/RawCullObjectAnalysisFeature.swift | Batch coordination, availability, cancellation, retry, private capture, and stage timings. |
ObjectAnalysis/ObjectConceptDiscovery.swift | Automatic prompt, concept validation, and Specific Concepts parsing. |
ObjectAnalysis/ObjectInstanceDeduplicator.swift | Filters weak masks, merges near-identical masks across concepts, and assigns board IDs. |
ObjectAnalysis/ObjectReviewBoardRenderer.swift | Renders the 2,048-pixel overview and numbered, outlined crops for Qwen; fails if a numbered crop cannot be prepared. |
ObjectAnalysis/ObjectJSONEnvelope.swift and ObjectAnalysis/ObjectAnalysisResponseDecoder.swift | Recover one JSON object from a wrapper, then validate fields, confidence, list limits, and board IDs. |
ObjectAnalysis/ObjectAnalysisModels.swift | Mode, instance, assessment, progress, timing, and result types. |
ObjectAnalysis/ObjectMaskOutlineRenderer.swift | Detail-view contour from a stored grayscale instance mask. |
Views/AIAnalysis/ObjectAnalysisView.swift | Controls, status table, numbered overlays, crop, per-object detail, and retry. |
The object-set workflow uses PhotoAIKit’s ObjectSegmentationService,
ObjectMaskMemoryStore, optional ObjectMaskDiskStore, and SAM 3
ObjectInstanceSegmenting contract. Its cache is separate from the Deep
Review subject-mask cache.
Qwen
| Source | Responsibility |
|---|
QwenInferenceRuntime.swift | Actor-owned Qwen provider validation, lazy vision-language model loading, session creation, prompt construction, response decoding, and invalidation. |
RawCullQwenAnalysisFeature.swift | Main-actor batch operation, image loading, progress, per-file failure isolation, result retention, and cancellation. |
QwenPhotoAssessment.swift | Structured response schema, validation, free-form fallback, aggregate score, and result types. |
Persistence and burst consumption
PerFileAnalysisArtifactStore persists descriptor-bearing similarity artifacts.
The burst-analysis files consume those artifacts, build groups, cache results,
and reject incompatible cache data. They are not model runtimes themselves.
This separation is important: a CLIP model produces an embedding; RawCull’s
burst policy decides what that embedding means for grouping and culling.
The Model Inventory Shipped by RawCull
The authoritative inventory is
RawCullAIModelDownloadCatalog.prepared.
production filters that inventory through code-only inclusion switches.
| Production model | Asset-pack ID | Installed model path inside pack | Download size | Installed size |
|---|
| DataComp CLIP, ViT-B/32 at 256 px | rawcull-clip-datacomp | Models/CLIP-DataComp | 282,967,354 bytes | 307,800,172 bytes |
| Meta SAM 3 | rawcull-sam3 | Models/SAM3 | 1,542,689,931 bytes | 1,667,570,378 bytes |
| Qwen3-VL-2B-Instruct | rawcull-qwen3-vl-2b | Models/Qwen/qwen3_vl_2b | 3,754,599,603 bytes | 5,395,195,663 bytes |
OpenAI CLIP exists in RawCullCLIPModel, has a resource manager, and can be
selected by the runtime, but includeOpenAICLIP is currently false.
DataComp CLIP, SAM 3, and Qwen download are enabled. SAM 3 requires explicit
acceptance of its bundled, hash-verified licence before download; the DataComp
and Qwen licences do not require an extra acceptance action.
How AI Modules Are Instantiated
The application has two stable roots: RawCullViewModel for general app state
and RawCullIntelligenceRuntime for AI-facing state. They are created once in
RawCullApp.init()
and retained in SwiftUI @State.
flowchart TD
App["RawCullApp.init()"] --> State["RawCullApplicationState.live()"]
State --> Models["RawCullAIModelRuntime"]
State --> Downloads["RawCullAIModelDownloadsModel"]
State --> QwenFeature["RawCullQwenAnalysisFeature"]
State --> Objects["RawCullObjectAnalysisFeature"]
State --> DeepFeature["DeepAIReviewFeature"]
State --> Settings["RawCullAISettingsModel"]
State --> Scoring["SimilarityScoringModel"]
Scoring --> Similarity["RawCullSimilarityFeature"]
Scoring --> Semantic["RawCullSemanticSearchFeature"]
DeepFeature --> Controller["DeepAIReviewController"]
State --> VM["RawCullViewModel"]
State --> Runtime["RawCullIntelligenceRuntime"]
Runtime --> Models
Runtime --> QwenFeature
Runtime --> Objects
Runtime --> Similarity
Runtime --> Semantic
Runtime --> Controller
Runtime --> SettingsThe exact construction order in RawCullApplicationState.make is significant:
RawCullAIModelRuntime is supplied by live(). Its initializer creates
resource-manager actors for SAM 3 and both CLIP choices, one
QwenInferenceRuntime, the Vision provider/service, mask stores, and a
placeholder unavailable segmentation pipeline.RawCullAIModelDownloadsModel is created with the production catalog and
application paths.RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature receive the
same Qwen inference actor. Objects also receives its memory and optional
disk instance-mask stores.DeepAIReviewFeature starts with the model runtime’s current segmentation
capability. It is bound back to the model runtime so a later SAM provider can
install a real pipeline without replacing the feature.RawCullAISettingsModel receives the model runtime, downloads model, Qwen
feature, preferences store, and saved-evidence scanner.- Settings produces a synchronous initial configuration. Before asynchronous
validation finishes this normally selects the Vision fallback.
- A single
SimilarityScoringModel is created. Both similarity and semantic
search share this same artifact/state owner. - Stable feature and controller objects are created around those models.
RawCullViewModel receives the exact same feature objects.RawCullIntelligenceRuntime retains the graph and binds the narrow weak
application contexts.- Settings binds its weak configuration consumer and immediately publishes
the first revision.
Debug assertions verify identity sharing. These checks are not cosmetic: a
second Qwen inference actor, scoring model, Objects feature, or Deep Review feature would split
model state, tasks, caches, and UI observation.
The first asynchronous validation begins from the main view’s .task:
.task {
await intelligenceRuntime.settingsModel.refresh()
}
refresh() asks the downloads model for an installed-location snapshot. That
snapshot flows through settings to RawCullAIModelRuntime, which validates
Qwen and refreshes CLIP and SAM 3 capabilities. See
The RawCull AI Runtime
for the concrete PhotoAIKit provider handoff, feature wiring, lifetime, and
reconfiguration path. In particular, the download snapshot supplies URLs;
PhotoAIKit factories validate bundles and create typed providers; the model
runtime retains those providers; and Settings sends selected services in a
revisioned configuration to the stable intelligence runtime. Qwen follows its
own actor path and updates its existing analysis feature through model status.
Why Qwen has its own inference runtime
It may look simpler to put Qwen’s provider, loaded model, and inference methods
directly inside RawCullAIModelRuntime. The two types have different jobs,
however, and keeping those jobs separate makes their concurrency and lifetimes
clear.
Think of RawCullAIModelRuntime as the coordinator for the application’s model
room. It knows which resources are installed, validates capabilities, selects
the services that RawCull should expose, and publishes those choices on the
main actor. QwenInferenceRuntime, by contrast, is the specialist operating
one machine in that room. Its actor protects Qwen-specific mutable state: the
validated provider, the lazily loaded vision-language model, and the generation
counter used to reject work from a model that has since been removed or
replaced. It also owns Qwen-specific work such as creating sessions, building
prompts, running generation, and decoding responses.
This boundary matters because Qwen inference can suspend for comparatively
long operations such as loading the model and generating a response. Those
operations should be serialized by the Qwen actor without turning the
main-actor RawCullAIModelRuntime into the place where heavy inference runs.
It also keeps Qwen’s two-stage lifecycle—validate a lightweight provider now,
then load the heavy model only when it is first used—independent of the CLIP
and SAM 3 resource lifecycles.
Separate does not mean unrelated. RawCullAIModelRuntime.init creates one
QwenInferenceRuntime and retains it as qwenInference. During application
assembly, that exact instance is passed to RawCullQwenAnalysisFeature and
RawCullObjectAnalysisFeature.
Consequently, there is one owner of Qwen’s provider and loaded model, while the
model runtime remains the composition point that creates and coordinates the
application’s complete collection of AI backends. In short:
RawCullAIModelRuntime answers which AI capabilities are available and
how they fit into the application;QwenInferenceRuntime answers how one Qwen request is safely executed;
and- constructing the latter inside the former guarantees a single, shared Qwen
runtime rather than independent copies with competing model state.
Model Discovery, Validation, and Provider Construction
CLIP and SAM 3 use one
RawCullAIModelResourceManager<Provider> actor per model choice. The actor owns:
- caller-ordered fallback candidate URLs;
- the current managed asset-pack URL;
- the PhotoAIKit
ModelProviderFactory; - a lightweight file-metadata snapshot; and
- the cached capability/provider result.
Changing the managed URL clears the cache. load() prepends the managed URL to
any fallback candidates, snapshots every directory entry, and returns its cached
result if path, kind, size, modification date, and resolved symlink path have not
changed. When the snapshot changes, PhotoAIKit remains authoritative: it checks
metadata.json, the declared model asset, required tokenizer files, accepted
.aimodel/.aimodelc extensions, and the model fingerprint or manifest checksum.
Only then does the factory create the concrete provider.
The file-metadata snapshot is an optimization, not a security decision. It
decides when full validation may be reused; it does not replace PhotoAIKit’s
model-bundle validation.
Qwen has a separate actor because its lifecycle differs. Its validation builds
an immutable CoreAIQwenProvider; its much heavier
CoreAIVisionLanguageModel is created lazily on the first assessment and reused
across later LanguageModelSession values.
How CLIP Works in RawCull
CLIP places images and text in a shared vector space. RawCull uses that property
in three ways: image-to-image distance, text-to-image semantic search, and a
small closed-set subject-label pass that helps choose SAM prompts.
Provider construction and identity
PhotoAIKit’s
CoreAICLIPProvider
is an actor implementing image embedding, artifact generation/comparison, text
embedding, and image/text comparison. During initialization it:
- validates the supplied model bundle;
- derives a
ModelIdentity and asset fingerprint; - decodes model-specific preprocessing, tokenizer, function-name,
normalization, and configuration metadata; and
- exposes a
SimilarityBackendDescriptor containing all compatibility-critical
versions.
The descriptor is effectively the type identity of an embedding on disk. It
records backend, model fingerprint, representation, preprocessing,
normalization, and configuration versions. Image artifacts additionally record
vector dimensions, schema version, and a source fingerprint. Consequently,
RawCull does not compare a DataComp vector with an OpenAI vector, reuse an
artifact after preprocessing changes, or silently treat an edited source file
as unchanged.
Image preprocessing and inference
RawCullSimilarityImageDecoder first asks RawParserKitImageLoader for a
bounded thumbnail. If that fails it attempts an ImageIO thumbnail without
requesting a full fallback decode. The resulting CGImage is passed through
the model-specific preprocessing declared by the CLIP bundle. Current metadata
supports either the legacy stretch/bilinear path or shortest-side resize plus
square center crop with bicubic interpolation and configured RGB mean/standard
deviation.
The provider lazily loads Core AI functions and tokenizer resources. The image
function receives the prepared image tensor and, if required by the exported
graph, dummy text inputs. Its output is flattened, checked against the expected
dimension, wrapped in an ImageEmbedding, JSON-encoded, and stored in a
SimilarityArtifact. The backend descriptor records the bundle’s declared
normalization version; image/text comparison later verifies that the image
vector actually has approximately unit magnitude.
Indexing and recovery
RawCullCLIPSimilarityService uses PhotoAIKit’s bounded
SimilarityArtifactIndexer with concurrency limit 1. CLIP inference is kept
serial because the provider is actor-owned and model execution is resource
intensive. There is no per-file or whole-batch Vision substitution during a
CLIP pass.
For a non-finite vector, RawCullRecoveringCLIPArtifactProvider performs a
targeted sequence:
- reject the invalid output;
- retry once with the already-loaded provider;
- construct a fresh provider from the same validated model location;
- verify that the replacement descriptor is exactly the same; and
- retry once with the replacement.
Successful CLIP artifacts from other files are retained. A file that still
fails is reported and remains unindexed; it is not given a Vision artifact that
would make the batch heterogeneous. Decode failures and inference failures are
recorded separately for diagnostics.
Image similarity and burst grouping
For compatible normalized image embeddings, PhotoAIKit returns cosine distance.
SimilarityScoringModel owns the artifact dictionary and computes distances
from an anchor. It can apply a small RawCull-owned subject-label mismatch
penalty, then uses those distances as input to ranking and burst grouping.
This responsibility split is deliberate:
- CLIP defines the vector and mathematical comparison;
- PhotoAIKit defines artifact compatibility and indexing mechanics; and
- RawCull defines catalog admission, grouping thresholds, ordering, progress,
persistence, and culling policy.
Vision remains the runtime fallback when CLIP is disabled or cannot be
validated. Vision and CLIP artifacts are both descriptor-bearing, so switching
backends causes incompatible state to be rejected or rehydrated rather than
misinterpreted.
Semantic search
Semantic search never indexes missing images as a side effect. It operates only
on already-persisted, descriptor-compatible CLIP image artifacts:
flowchart LR
Query["Text query"] --> Tokens["CLIP tokenizer"]
Tokens --> TextModel["CLIP text function"]
TextModel --> TextVector["Validated normalized text vector"]
Images["Compatible cached image artifacts"] --> Compare["Dot product / cosine similarity"]
TextVector --> Compare
Compare --> Sort["Descending score with deterministic tie breaks"]RawCullCLIPSemanticSearchService trims and validates the query, filters image
artifacts by the complete backend descriptor, generates one transient text
embedding, and scores each compatible image. Because both vectors are
normalized, their dot product is cosine similarity. Valid scores lie in
-1...1; they are relative ranking values, not confidence percentages.
Sorting is deterministic: score descending, then original catalog order,
localized filename, and UUID. Individual malformed artifacts become per-file
failures rather than aborting every valid result. Text embeddings are scoped to
one search and are not persisted.
Deep Analysis: How SAM 3 and CLIP Work Together
The UI calls this mode SAM 3 + CLIP, but the two models are not fused and
CLIP does not calculate the final focus score. Their collaboration is a staged
pipeline:
- CLIP optionally supplies a coarse subject label from existing embeddings.
- That label selects an ordered set of text prompts for SAM 3.
- SAM 3 creates or retrieves the best acceptable subject mask.
- RawCull measures detail only inside that mask and recommends the strongest
candidate.
If CLIP semantic artifacts are unavailable, RawCull falls back to the existing
saliency label from normal sharpness analysis. If neither label exists, SAM 3
still receives the general subject prompt. Therefore SAM 3 is the required
model for Deep Review; CLIP enriches prompt selection when available.
1. Building the request
DeepAIReviewController.start(for:) asks RawCullViewModel for a stable group
context. deepAIReviewContext(for:) captures:
- a
BurstGroupSignature tied to the current catalog and exact member files; - existing burst ranks, falling back to input order;
- normal sharpness scores;
- a subject label;
- the normalized camera autofocus point; and
- the chosen sharpness source: embedded preview or RAW demosaic.
For the CLIP label pass, SimilarityScoringModel.classifySubjects runs six
literal queries against existing semantic artifacts: person, bird, deer,
animal, car, and landscape. Each file receives the label with its highest
cosine similarity. This pass does not alter the visible semantic-search result,
decode source images, or generate missing embeddings.
2. Candidate limiting and decoding
Candidates are sorted by current burst rank. Groups of 12 or fewer are analyzed
in full; larger groups analyze the first 8 candidates. Each candidate is then
decoded at a bounded size. Embedded-preview mode reuses RawCull’s similarity
decoder. RAW-demosaic mode uses CIRAWFilter, explicitly sets sharpness to 0,
detail to 0.6, contrast to 1, and exposure to 0, then downsizes to the configured
maximum before producing a CGImage.
3. Prompt selection
The preset and subject label determine ordered attempts:
| Preset/evidence | SAM prompt order |
|---|
| Full Subject | subject |
| Auto or Head/Face with bird/wildlife label | bird head, bird, subject |
| Auto or Head/Face with person/face label | face, person, subject |
| Auto or Head/Face with deer label | animal head, deer, animal, subject |
| Auto or Head/Face with generic animal label | animal head, animal, subject |
| No recognized label | subject |
The Head/Face preset is considered verified only when the first selected prompt
is bird head, animal head, or face; falling back to a broader mask is
reported as specificPromptNotFound.
4. SAM 3 inference
PhotoAIKit’s
CoreAISAM3Provider
is actor-isolated. It validates the bundle and lazily creates a
CoreAISegmentationEngine plus CLIP-compatible text tokenizer. The prompt text
is tokenized and sent with the bounded image to the Core AI segmenter. Runtime
parameters use a 0.5 mask threshold and at most 5 segments.
The response’s probability map is preferred. If absent, the provider unions
compatible returned segment masks. Probabilities are converted to a white RGBA
mask with a smooth alpha transition around the threshold. The result records
the prompt, confidence, model identity, input/output sizes, timing, resource,
and asset identity.
PhotoAIKit’s SegmentationService first checks memory and disk stores by a key
that includes source file identity, prompt, model identity, and maximum input
size. A missing mask is generated, resized back to the display image dimensions,
and saved to both stores. Model changes therefore do not accidentally reuse
masks made by a different SAM asset.
SubjectMaskSelector tries prompts in order, measuring coverage and quality for
each candidate. It stops at the first mask meeting the warning-or-better
threshold, otherwise retains the best attempt by quality, confidence, and then
coverage. Every attempt records cache miss, candidate quality/confidence, or
failure.
5. Subject-detail scoring
SubjectMaskFocusScorer converts the photograph to luminance using Rec. 709
weights and samples Laplacian-style edge energy. Only pixels with mask alpha
above 16 contribute to subject evidence. It computes:
- broad subject score — robust tail detail across the masked subject;
- local detail score — the strongest reliable cell in a 6 × 6 patch grid;
- fine detail score — micro-contrast within the mask;
- mask coverage — masked pixels divided by total pixels; and
- AF evidence — whether the normalized autofocus point lies inside the mask.
The final detail score is:
0.40 × broad subject detail
+ 0.40 × strongest local detail (or broad detail when local is unavailable)
+ 0.20 × fine detail
If global edge detail exceeds the subject score by the configured margin,
RawCull applies a 0.82 background-dominance multiplier and records a caution.
The scorer reports missing local patches, unusable masks, unavailable subject
detail, and other evidence limitations rather than manufacturing a score.
6. Recommendation and confidence
Candidates sort by deep score descending, with the earlier burst rank breaking
ties. The first finite score becomes the recommendation. Reasons record strong
subject detail, AF-inside-subject evidence, local detail evidence, and successful
prompt matching.
Confidence depends on the winning margin and evidence quality:
- high: at least a 12% lead, mask and local evidence present, no issues, and
no fallback prompt;
- medium: at least a 5% lead, or strong evidence obtained through a fallback
prompt; and
- low: all other cases.
The result is advisory. Deep Review stores results by group signature, publishes
progress after every candidate, and keeps completed candidate/mask evidence for
the analysis history and zoom outline. It does not silently change ratings or
apply a culling decision.
How Qwen Analysis Works
Qwen is not part of CLIP similarity or SAM segmentation. It is a separate local
vision-language tool selected in the AI Analysis view.
Validation and lazy loading
QwenInferenceRuntime is an actor with three pieces of state: an optional
CoreAIQwenProvider, an optional loaded CoreAIVisionLanguageModel, and a
generation counter. Validation asks PhotoAIKit’s Qwen factory to inspect the
bundle. The provider checks required tokenizer resources, metadata kind, Qwen
identity, vocabulary/context values, and VLM-specific embedding and vision
assets. RawCull additionally rejects a valid text-only Qwen bundle because
photo assessment requires modality .vision.
Successful validation stores the lightweight provider and clears any previously
loaded model. The first call to assess constructs the heavy vision-language
model asynchronously; later calls reuse it. A generation captured before
loading prevents a model removed or replaced during the await from becoming
active afterward. clear() advances the generation and releases both provider
and loaded model.
Per-image request
The feature processes pending files sequentially. It requests a thumbnail up to
2048 pixels, then calls assess(criteria:image:). The runtime creates a new
Foundation Models LanguageModelSession around the reused model, attaches the
CGImage, and requests at most 512 response tokens.
The prompt includes the user’s criteria and asks for exactly one JSON object
when the request is a photo assessment. The schema contains:
- subject description;
- composition, exposure, and subject-visibility scores from 1 through 5;
- optional
eyesOpen; - up to four problems and strengths; and
- confidence from 0 through 1.
The instruction explicitly says to use visible evidence only. For a request
that does not fit the assessment schema, Qwen may answer in ordinary text.
Response handling
QwenModelResponse.decode trims the response. It first attempts to extract the
outermost JSON object and decode QwenPhotoAssessment. Score ranges and
confidence are validated. If that structured decode does not succeed but the
response is nonempty, RawCull retains it as a free-form result.
The structured overallScore is a RawCull presentation value, not a model
output:
0.50 × composition
+ 0.20 × exposure
+ 0.30 × subject visibility
Each component is first divided by 5. The batch continues after an individual
file fails, and successful or failed results replace earlier results for the
same file. Already analyzed files are skipped on the next run. Cancellation and
model removal advance operation state so obsolete work cannot publish as a
current result.
Qwen results currently live in the feature’s in-memory results array; unlike
CLIP artifacts and SAM masks, this implementation does not persist them across
application sessions.
Objects: instance-level SAM 3 and Qwen analysis
Objects is a third AI Analysis tool beside SAM 3 + CLIP and standalone
Qwen. It accepts selected Grid photos or tagged photos. It requires
installed, validated SAM 3 and vision-capable Qwen. CLIP embeddings and Deep
Review’s single-subject score do not feed this workflow.
End-to-end stages and ownership
- The stable, main-actor object feature checks that both model services are
available and snapshots the concept mode and photographic criteria.
- It loads one bounded RAW or JPEG thumbnail, at most 4,320 pixels on its
longest side, and processes files sequentially.
- Automatic mode asks Qwen for visible object concepts; Specific Concepts
parses the user’s comma-separated noun phrases.
- PhotoAIKit’s object service asks SAM 3 for up to eight instances per concept
and checks its separate object-mask caches.
- RawCull filters weak/invalid masks, merges near-duplicate regions across
concepts, and assigns board-local IDs 1 through 8.
- A deterministic 2,048-pixel board shows the original overview above
numbered, outlined object crops. Qwen assesses that one photograph.
- RawCull extracts one complete JSON object and validates the schema, every
expected board ID, values, list caps, and finite confidence. The UI shows
per-object findings or a visible, retryable assessment problem.
The availability state distinguishes checking, ready, SAM 3 unavailable,
Qwen unavailable, and both unavailable. Settings installs the current
segmentation service and Qwen status into the existing feature after
validation. A changed service or status cancels active work. Switching AI
tools or input source also cancels an active batch. Completion, result
replacement, and cancellation are generation-gated.
Concept discovery and manual mode
Automatic asks Qwen for zero to six short concept entries. Each entry has
query, displayName, and reason. Query must be a concrete, visible, whole-object
noun phrase suitable for SAM 3. SegmentationConcept validates it; normalized
duplicates collapse. An empty set or invalid JSON is an actionable discovery
failure, with no guessed fallback concepts. The discovery request cap is 384
output tokens.
Specific Concepts bypasses discovery. The user enters up to six
comma-separated queries such as bird, person. Empty or invalid entries fail
before segmentation; duplicate normalized queries collapse. Both modes can
add photographic criteria to the final Qwen request.
SAM 3 instances, filtering, and caches
ObjectSegmentationService uses source file identity, concept, SAM 3 model
identity, the 4,320-pixel input limit, and the eight-instance limit in its
cache key. It checks the object-mask memory and optional disk store before
inference, bounds the image, invokes CoreAISAM3Provider.segmentInstances,
resizes masks to display dimensions, and saves the typed result. This is
independent of Deep Review’s SegmentationService and SubjectMaskSelector,
which choose one subject mask from ordered prompt attempts.
RawCull counts the raw SAM 3 candidates. Its deduplicator rejects nonfinite
or below-0.5 mask scores, invalid boxes, masks with fewer than 64 of
256-by-256 sampled pixels, and masks covering at least 95% of the sample.
Candidates sort by score and geometry. Overlapping masks with similar area
merge when mask intersection-over-union reaches 0.85 or smaller-mask
containment reaches 0.90; a second concept becomes an alias of the retained
object. At most eight objects remain. Their board IDs are strings and local to
that analysis result. A SAM 3 mask score describes segmentation quality; it
is distinct from Qwen’s assessment confidence.
An empty retained set is a successful No Matching Objects result and does
not call Qwen for a board. Genuine no-match photos still need broader
validation.
Board geometry and the Qwen contract
ObjectReviewBoardRenderer creates one 2,048-by-2,048 image. Its top half is
an aspect-fit overview of the source photo. Its bottom half contains up to
eight padded crops in two rows of four. Each crop and yellow mask outline are
aspect-fit. A separate dark header carries a large white number without
covering the subject. The renderer explicitly converts normalized
bottom-left box coordinates to top-left CGImage crop coordinates.
If the source or mask crop for a retained object cannot be made, rendering
throws reviewBoardUnavailable. The feature records the per-photo failure
before asking Qwen to assess a board; it does not submit a board with a
numbered ID whose crop was silently omitted.
The prompt tells Qwen that the overview and crops repeat views of one
photograph, each board ID denotes a different physical subject, and the
objects array must contain exactly one entry for each ID. Descriptions should
use the matching numbered crop; relationships should use the overview.
The assessment request cap is 1,024 output tokens. The requested result has
an optional scene summary; per-object concept, description, visibility, focus,
expression, obstructions, strengths, problems, and confidence; and photo-level
relationships, strengths, problems, preferred IDs, and confidence.
ObjectJSONEnvelope extracts one complete balanced JSON object from a
recoverable Markdown or prose wrapper, respecting quoted braces. It rejects
incomplete JSON and multiple objects. The decoder accepts only the narrow
variations observed from the packaged model: an omitted imageSummary, a
whole-number board ID normalized to a string, and one short text value in an
object list field (the exact string “none” becomes an empty list). It still
requires every board ID exactly once, rejects unknown or duplicate IDs and
preferred IDs, enforces list limits and enum values, and requires finite
confidence in 0…1. Free-form prose is not upgraded to a structured judgment.
If Qwen fails this boundary, the detail view retains its response and shows
a specific assessment error. The results table says Assessment needs
retry. Retry reuses successful SAM 3 masks only when source size/date, Qwen
model name, concept mode and queries, and cache keys still match. Otherwise
it segments again. The table labels Qwen confidence; the detail view labels
SAM 3 mask score separately and shows numbered boxes, cached-mask outlines,
the selected crop, and its assessment.
The object count in the table is the number of retained SAM 3 matches for the
chosen concepts, after filtering and deduplication. It is not a count of all
subjects in the photograph. A Complete row means the response passed the
schema and board-ID checks; the photographer still needs to compare its
descriptions, count claims, and confidence with the source image. The detail
panel shows the assessment for the currently selected object, not all object
descriptions at once.
September 24, 2026 packaged-model check
Eight supplied puffin ARW files were processed in both modes with local
qwen3_vl_2b and sam3_float16.aimodel. Automatic discovered puffin for all
eight; manual mode used bird. All 16 runs produced structured assessments
with the expected board ID set. SAM 3 returned seven or eight raw candidates
per run, filtered to one retained bird on six photos and two birds on two.
Warm concept discovery took 2.98–3.56 seconds, segmentation 6.59–6.99
seconds, board rendering 0.028–0.044 seconds, and final Qwen assessment
10.21–20.14 seconds. The sequential probe took 371.5 seconds; process
peak resident memory was 9.47 GiB. The first Automatic run included model
startup and took 45.1 seconds overall.
These measurements came from an in-process macOS feature/model test, not a
release build, clean install, or TestFlight run. Separate two-bird visual
checks found a swapped flying/perched description, nearly identical
descriptions for differently facing birds, and a response claiming three
birds where the photograph had two. Those errors occurred with schema-valid
output and Qwen-reported confidence of 0.95–1.00. Valid JSON and IDs verify
response shape, not visual accuracy or calibrated confidence. Mixed
categories, touching/overlapping and tiny subjects, genuine no-match cases,
RAW/JPEG parity, model removal, large tagged batches, and lifecycle checks
remain open. The per-photo table and gate status are in the
RawCull implementation notes.
September 24–26, 2026 in-app observations
In-app Automatic runs showed all eight puffin photos and all eight photos in a
mixed-subject batch as Complete. The mixed batch included landscape, deer,
muskox, horse, bird, rabbit, and puffin photographs. These screenshots show the
workflow operating in the app on those selections, but do not identify the
installed model-pack fingerprints or establish a clean-install TestFlight run.
Two visual checks remain unresolved. In _DSC3028.ARW, clicking the two
numbered puffins showed crops and descriptions associated with opposite birds;
the source of the mismatch, whether Qwen’s board grounding or the UI’s
object-to-crop association, has not been established. In _DSC3031.ARW, SAM 3
retained two puffins and Qwen’s object details described two positions, while
its image summary claimed a third puffin on the ground. That Complete result
reported 95% Qwen confidence. The invented third bird is a prose error, not a
third SAM 3 instance or a missing board ID.
Objects is being treated as an advisory test feature while users report issues.
The two photographs are regression cases for visual grounding and object
mapping. Before treating the broader validation as complete, the signed build
and hosted model packs still need a clean-install run covering launch, analyze,
cancel, remove, and reinstall, with build, model identities, macOS version, and
memory recorded. See the in-app validation and follow-up
and September 26 observation.
Private diagnostics
The feature records raw/retained counts, model identities, and separate
concept-discovery, segmentation, board-rendering, and assessment durations.
An explicit RAWCULL_OBJECT_CAPTURE_DIR environment variable enables private
response capture with file name/ID, stage, model, requested token cap, and
response character count. The directory must have 0700 permissions and new
files use 0600. Normal logs do not include full images, prompts, or Qwen
responses. The runtime exposes neither finish reason nor generated-token
count, so response length and visible truncation must be inspected directly.
Remove private capture files after diagnosis.
Capability and Failure Behavior
The settings UI distinguishes these states instead of reducing them to one
Boolean:
checking: a location exists or is being resolved, but validation is not
complete;available: validation and provider construction succeeded;missing: expected resources are absent;invalid: a resource exists but metadata, checksum, files, modality, or
provider construction failed; andunavailable: a feature cannot be offered for another explicit reason.
Semantic-search readiness is separate from generic CLIP readiness. Image
similarity can always fall back to Vision; text search requires a validated
provider that implements the text/image contracts. Deep Review can keep its
stable controller while its service is temporarily nil. Qwen publishes its
own status and cancels an active batch if that status becomes unavailable.
Persistence Boundaries
| Data | Lifetime/location | Compatibility protection |
|---|
| CLIP or Vision similarity artifacts | Per-file analysis artifact store and burst cache | Full backend descriptor plus source fingerprint and schema version |
| CLIP text query embedding | One search call | Never persisted |
| SAM 3 masks | Memory store plus Caches/no.blogspot.RawCull/SAM3Masks when disk-store construction succeeds | Source identity, prompt, model identity, and max input side |
| Deep Review recommendations | In-memory feature dictionary keyed by BurstGroupSignature | Exact group signature; reset/cancellation generation |
| Qwen results | In-memory feature array keyed by file UUID | Current batch generation; no cross-launch persistence |
| Objects masks | Separate object-mask memory store and optional ObjectMaskDiskStore | Source identity, concept, SAM 3 model identity, 4,320-pixel input limit, and eight-instance limit |
| Objects assessments and timings | In-memory feature results keyed by file UUID | Batch generation and board-ID validation; no cross-launch assessment persistence |
| User model selections | UserDefaults | Inclusion lists sanitize choices no longer shipped |
| Model assets | Managed Background Assets locations | Catalog ID, model bundle validation, and asset fingerprint/checksum |
Practical Trace Points
When debugging a model problem, follow the layer that owns the decision:
- Asset not present or licence blocked: model download catalog, downloads
model, and download service.
- Bundle present but invalid:
RawCullAIModelResourceManager and
PhotoAIKit ModelBundleResolver. - Provider validates but feature stays on Vision: settings snapshot,
RawCullAIModelRuntime.similarityService, and runtime configuration identity. - Some CLIP images fail: decoder/inference failure report and finite-vector
recovery in
RawCullCLIPSimilarityService. - Semantic search has no candidates: semantic artifact hydration and exact
descriptor compatibility.
- SAM mask is missing or poor: prompt attempts, mask cache key, geometry,
quality, and segmentation diagnostics.
- Deep score looks unexpected: inspect broad/local/fine evidence, mask
coverage, AF inclusion, and background-dominance caution.
- Qwen is available but a batch fails: distinguish thumbnail decoding,
lazy model load, session response, empty response, and per-file result decode.
- Objects fails before SAM 3: inspect concept discovery and its exact JSON
or concept-validation error; Specific Concepts isolates that boundary.
- Objects has masks but no structured judgment: inspect the assessment
error, rendered board, and board-ID set. Retry may reuse cached masks.
- Objects text disagrees with the photo: compare the source, numbered crop,
outline, and description. High model-reported confidence does not settle a
grounding error.
The central architectural rule is that model runtimes create typed evidence;
RawCull’s feature and policy layers decide how that evidence affects ranking,
grouping, presentation, and user actions.
10.4 - The RawCull AI Runtime
How RawCull owns and refreshes local CLIP, SAM 3, Qwen, Vision, Deep Review, and Objects runtimes.
The RawCull AI Runtime
RawCull’s AI runtime is the long-lived object graph that connects downloaded
model assets to stable application features. It is not one model, one thread,
or a background daemon. It is a set of objects with deliberately different
lifetimes and actor-isolation rules.
The current implementation has two runtime layers:
| Runtime | Primary responsibility |
|---|
RawCullAIModelRuntime | Own concrete provider/resource lifecycles: CLIP, SAM 3, Qwen, Vision, model capability snapshots, separate subject/object mask stores, and segmentation-service installation. |
RawCullIntelligenceRuntime | Own stable similarity, semantic search, Deep Review, Qwen, and Objects feature lifetimes; apply complete, revisioned similarity/semantic/segmentation settings decisions without rebuilding the graph. |
That distinction replaces the older, broader RawCullAIIntegration shape. The
authoritative sources are
RawCullAIModelRuntime.swift
and
RawCullIntelligenceRuntime.swift.
For model algorithms and data products, see
AI Models in RawCull.
Runtime Topology
flowchart TD
App["RawCullApp"] --> AppState["RawCullApplicationState"]
AppState --> VM["RawCullViewModel"]
AppState --> Runtime["RawCullIntelligenceRuntime"]
Runtime --> ModelRuntime["RawCullAIModelRuntime"]
Runtime --> Settings["RawCullAISettingsModel"]
Runtime --> Downloads["RawCullAIModelDownloadsModel"]
Runtime --> Similarity["RawCullSimilarityFeature"]
Runtime --> Semantic["RawCullSemanticSearchFeature"]
Runtime --> Review["DeepAIReviewController"]
Runtime --> Qwen["RawCullQwenAnalysisFeature"]
Runtime --> Objects["RawCullObjectAnalysisFeature"]
ModelRuntime --> CLIP["CLIP resource managers/providers"]
ModelRuntime --> SAM["SAM 3 resource manager/provider"]
ModelRuntime --> QwenActor["QwenInferenceRuntime actor"]
ModelRuntime --> Vision["Vision fallback"]
ModelRuntime --> Masks["Mask repository/stores/selector"]
ModelRuntime --> ObjectMasks["Separate object-mask stores and instance service"]
Similarity --> SharedModel["SimilarityScoringModel"]
Semantic --> SharedModel
Review --> DeepFeature["DeepAIReviewFeature"]
Qwen --> QwenActor
Objects --> QwenActor
Objects --> ObjectMasksThe arrows above mix ownership and collaboration. The precise ownership rules
are discussed below; notably, callbacks from a child toward an owner are weak.
What Each Layer Owns
RawCullAIModelRuntime
The model runtime is @MainActor because model selection and provider
installation must be coordinated with observable feature state. Heavy work is
still isolated elsewhere: its resource managers and Qwen runtime are actors,
and PhotoAIKit’s CLIP and SAM 3 providers are actor-owned.
It owns:
- application AI paths;
- three
RawCullAIModelResourceManager actors: SAM 3, DataComp CLIP, and
OpenAI CLIP; - a single
QwenInferenceServing actor; - the always-available
VisionFeaturePrintBackend and Vision similarity
service; - dictionaries of validated CLIP and segmentation providers;
- resolved CLIP model locations used to create replacement providers;
- memory and optional disk subject-mask stores;
- separate memory and optional disk object-mask stores;
- an optional
ObjectSegmentationService built from the validated SAM 3
provider and those object stores; - the current
SubjectMaskRepository, SegmentationService, and
SubjectMaskSelector; - the selected segmentation model and active model identity; and
- the latest
RawCullAICapabilities snapshot.
Views do not traverse this object. They receive the focused feature surfaces
from RawCullIntelligenceRuntime.
RawCullIntelligenceRuntime
The intelligence runtime owns the objects whose identities must remain stable:
let modelRuntime: RawCullAIModelRuntime
let similarityFeature: RawCullSimilarityFeature
let semanticSearchFeature: RawCullSemanticSearchFeature
let deepAIReviewController: DeepAIReviewController
let qwenAnalysisFeature: RawCullQwenAnalysisFeature
let objectAnalysisFeature: RawCullObjectAnalysisFeature
let settingsModel: RawCullAISettingsModel
let modelDownloadsModel: RawCullAIModelDownloadsModel
It also records the last accepted configuration revision and identity. Its
single mutation entry point is apply(configuration:).
RawCullViewModel
The main view model owns product and catalog policy, not model runtimes. It
answers questions such as which files are selected, what the current catalog
identity is, what burst ranks and sharpness scores exist, and whether an AI
operation conflicts with other work. Narrow protocols expose only the pieces
the AI features need.
Construction: From App Launch to a Live Graph
RawCullApp.init()
calls RawCullApplicationState.live(). SwiftUI retains the returned view model
and intelligence runtime in separate @State properties:
RawCullApp
├─ strong → RawCullViewModel
└─ strong → RawCullIntelligenceRuntime
live() constructs a default RawCullAIModelRuntime, then delegates to the
injectable RawCullApplicationState.make(...). The factory accepts stores,
preferences, scanners, download catalog/coordinator, and version information as
parameters so tests can build the same graph with deterministic substitutes.
Phase 1: initialize the model runtime
RawCullAIModelRuntime.init performs only synchronous, bounded setup:
- Store
RawCullAIPaths and the Qwen inference actor. - Create SAM 3 and CLIP resource-manager actors with their PhotoAIKit
factories. Managed URLs are initially unset.
- Create one Vision provider and wrap it in
RawCullVisionSimilarityService. - Create
SubjectMaskMemoryStore. - Attempt to create
SubjectMaskDiskStore at
Caches/no.blogspot.RawCull/SAM3Masks. Failure is represented as a
capability state; it does not prevent the application from launching.
Create a separate object-mask memory store and try an object disk store at
the configured object-mask cache path. A disk-store failure leaves Objects
with its memory store. - Build the list of usable stores: memory always, disk when construction
succeeded.
- Create an
UnavailableSegmentationProvider, repository, segmentation
service, and selector. This placeholder gives the graph a complete shape
before SAM validation. - Publish an initial capability snapshot: Vision available, model resources
checking, and mask-storage status known.
No CLIP, SAM 3, or Qwen model engine is loaded in this initializer.
Phase 2: assemble stable features
RawCullApplicationState.make then performs the following order:
- Create
RawCullAIModelDownloadsModel from runtime paths and the production
model catalog. - Create
RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature with
the same modelRuntime.qwenInference. Pass the object feature the exact
memory and optional disk object-mask stores owned by the model runtime. - Create
DeepAIReviewFeature with the initial mask-generation capability. - Bind that exact feature to the model runtime. The runtime installs an actual
pipeline later when a segmentation provider becomes available.
- Create
RawCullAISettingsModel with model runtime, downloads model, Qwen
feature, user defaults, and saved-burst-evidence scan. - Ask settings for a synchronous revision-0 configuration.
- Create one
SimilarityScoringModel from the selected similarity service,
semantic capability/service, and persistent artifact store. - Wrap it in
RawCullSimilarityFeature and
RawCullSemanticSearchFeature. Both wrappers refer to the same scoring model. - Wrap the Deep Review feature in
DeepAIReviewController. - Create
RawCullViewModel with those exact feature/controller instances. - Create
RawCullIntelligenceRuntime and bind the similarity feature’s weak
application context. - Bind semantic search to the view model and settings to the runtime.
The final settings binding immediately publishes the first configuration. This
happens only after every receiver exists.
Phase 3: verify graph identity
Debug assertions verify that:
- view model and runtime share the same similarity feature;
- semantic search and similarity share the same scoring model/feature identity;
- view model and runtime share the same semantic and Deep Review objects;
- controller and model runtime refer to the same Deep Review feature;
- runtime and Qwen feature share the same Qwen inference actor;
- runtime and Objects feature share that same Qwen inference actor;
- settings, runtime, and Qwen feature share the intended model runtime and
inference actor; and
- settings and runtime expose the same downloads model.
These are architectural invariants. Two equivalent-looking instances would not
share task handles, progress, caches, result dictionaries, generation counters,
or SwiftUI observation.
Development Handoff: PhotoAIKit Objects to the Runtime
PhotoAIKit supplies provider factories, typed contracts, workflows, and stores.
RawCull creates those objects and decides when they become usable. There is no
package callback that injects providers into RawCullIntelligenceRuntime:
RawCullAIModelRuntime owns the provider handoff, while Settings delivers
selected services to the stable feature objects.
Declare and construct the package boundary
RawCullAIModelRuntime.swift imports CoreAICLIPBackend,
CoreAISAM3Backend, PhotoAIContracts, PhotoAIStorage,
PhotoAIWorkflows, and VisionFeaturePrintBackend. At construction it passes
CoreAICLIPProvider.factory and CoreAISAM3Provider.factory to separate
RawCullAIModelResourceManager actors. It also creates the Vision provider,
SubjectMaskMemoryStore, optional SubjectMaskDiskStore, and an unavailable
segmentation provider. The latter lets SegmentationService,
SubjectMaskRepository, and SubjectMaskSelector exist before a SAM bundle is
validated. Qwen’s CoreAIQwenProvider.factory is used by the separate
QwenInferenceRuntime actor.
These are the objects crossing the package boundary:
| PhotoAIKit object or contract | RawCull receiver and use |
|---|
CoreAICLIPProvider and SimilarityBackendDescriptor | Model runtime retains the validated provider; similarity and semantic services wrap it. The descriptor identifies compatible artifacts. |
CoreAISAM3Provider as SubjectSegmenting, with ModelIdentity | Model runtime installs it in a segmentation service and selector, then gives Deep Review a pipeline. Identity determines when the pipeline must be rebuilt. |
The same CoreAISAM3Provider as ObjectInstanceSegmenting | Model runtime builds a separate ObjectSegmentationService with object-mask stores and gives it to the existing Objects feature through Settings. |
VisionFeaturePrintBackend | Model runtime wraps it in the always-ready Vision similarity service. |
SubjectMaskMemoryStore and optional SubjectMaskDiskStore | Repository and segmentation service share the stores; the disk store also supports Deep Review mask loading. |
ObjectMaskMemoryStore and optional ObjectMaskDiskStore | A separate instance-mask cache namespace, shared by object segmentation, detail outline loading, and assessment retry. |
CoreAIQwenProvider | Qwen inference actor retains the validated provider and loads its vision-language model on first use; the Qwen feature holds that same actor. |
Enable installed models and hand off providers
The development path from a model download to a live feature is:
Background Assets snapshot (complete model-ID → URL map)
→ Downloads model → Settings.applyManagedModelLocations
→ RawCullAIModelRuntime.applyManagedModelLocations
→ resource-manager actors / QwenInferenceRuntime
→ PhotoAIKit capability check and provider construction
→ RawCullAIModelRuntime.refreshCapabilities
→ Settings.configurationSnapshot → RawCullIntelligenceRuntime.apply
→ existing feature objects
The model runtime supplies managed URLs to each CLIP and SAM resource manager.
Each actor checks a metadata snapshot, asks its PhotoAIKit factory to validate
the candidate bundle, and constructs a provider only for an available resource.
refreshCapabilities() loads those actors concurrently, then stores the
validated providers, their resolved CLIP URLs, and a capability snapshot on the
main actor. If SAM 3 validated, it also constructs the object instance service
with the separate object-mask stores. Qwen instead validates or clears its actor in
applyManagedModelLocations; Settings passes its status to the already-created
Qwen feature. Validation does not eagerly load Qwen’s vision-language model.
The provider reaches a feature through one of three paths:
- For image similarity, Settings calls
modelRuntime.similarityService(prefersCLIP:clipModel:). It wraps the selected
validated CLIP provider in RawCullCLIPSimilarityService, or uses the Vision
service. The CLIP service also receives the resolved bundle URL through a
replacement-provider factory for finite-vector recovery. - For semantic search, Settings calls
modelRuntime.semanticSearchService(clipModel:). The same validated CLIP
provider backs RawCullCLIPSemanticSearchService; a missing provider yields
no semantic service. configurationSnapshot carries the selected capability
and service to RawCullIntelligenceRuntime.apply(configuration:), which
updates the shared SimilarityScoringModel through its stable feature. - For Deep Review,
refreshCapabilities() activates the selected SAM
provider. The model runtime rebuilds the repository, segmentation service,
and selector only when ModelIdentity changes, then calls
DeepAIReviewFeature.install with a pipeline, optional disk-mask loader, and
availability. The existing controller and feature keep their identities. - For Objects, Settings installs
modelRuntime.objectSegmentation and the
current Qwen status into RawCullObjectAnalysisFeature after validation.
The feature stays at the same identity and becomes ready only when both
services are available. It uses the same Qwen actor as standalone Qwen and
the same SAM 3 provider as Deep Review, through a different workflow/cache.
The runtime configuration carries selected services and descriptor-based
identity, not an unvalidated model URL. Its revision prevents an older Settings
decision from overwriting a newer one. When adding a PhotoAIKit backend, wire
its factory and managed location into the model runtime, translate its
capability and provider result, then expose it through the appropriate stable
feature or configuration path. Keep model-specific inference inside the
provider or actor and application policy inside RawCull.
Startup Refresh and Installed-Model Activation
The graph is usable immediately with Vision while disk checks happen later.
When RawCullMainView appears, this task starts the real refresh:
.task {
await intelligenceRuntime.settingsModel.refresh()
}
The call path is:
sequenceDiagram
participant View as RawCullMainView
participant Settings as RawCullAISettingsModel
participant Downloads as RawCullAIModelDownloadsModel
participant Models as RawCullAIModelRuntime
participant Resource as Resource-manager actors
participant Runtime as RawCullIntelligenceRuntime
View->>Settings: refresh()
Settings->>Downloads: refresh()
Downloads-->>Settings: applyManagedModelLocations(snapshot)
Settings->>Models: applyManagedModelLocations(snapshot)
Models->>Models: validate or clear Qwen
Models-->>Settings: Qwen status
par model validation
Settings->>Models: refreshCapabilities()
Models->>Resource: load SAM 3 and both CLIP choices
and saved evidence
Settings->>Settings: scan burst caches
end
Models-->>Settings: capabilities
Settings->>Runtime: apply(revisioned configuration)
Runtime-->>Settings: current capabilitiesOne complete location snapshot
applyManagedModelLocations(_:) is the only activation path for a complete set
of installed locations. It gives the current SAM and CLIP URLs to their resource
managers. A missing Qwen URL calls qwenInference.clear(); a present URL is
standardized and validated.
Using a complete snapshot avoids a transient mixture such as “new CLIP, old
SAM, removed Qwen still active.” Every invocation describes one model-install
state.
Generation-gated refresh
Settings increments refreshGeneration before beginning work. The generation
is checked after Qwen validation and after the concurrent capability/evidence
work. A later refresh therefore supersedes an earlier one; the earlier result
cannot publish merely because its disk work finished last.
The defer that clears isScanningSavedBurstData also checks the generation,
so an obsolete refresh cannot hide the current refresh’s progress indicator.
Concurrent capability validation
RawCullAIModelRuntime.refreshCapabilities() starts SAM 3, DataComp CLIP, and
OpenAI CLIP loads with async let. Each resource-manager actor computes a
lightweight recursive metadata snapshot. When unchanged, the previous validated
capability/provider result is reused. When changed, PhotoAIKit validates the
bundle and constructs a provider.
After all three complete, the main-actor runtime atomically replaces its
provider dictionaries and resolved location dictionaries, translates package
statuses into RawCull statuses, builds semantic-search readiness separately,
stores one new capability snapshot, and activates the selected segmentation
provider.
Capability State Is More Than “Loaded”
RawCullAICapabilityStatus distinguishes:
| State | Meaning |
|---|
checking(expectedLocations:) | Validation is pending. |
available(location:) | Bundle validation and provider construction succeeded, or an always-available service such as Vision is ready. |
missing(expectedLocations:) | No candidate bundle was found. |
invalid(location:reason:) | A candidate exists but validation or provider construction failed. |
unavailable(reason:) | The runtime cannot offer the capability for another explicit reason. |
CLIP model availability and semantic-search readiness are separate. A provider
must expose the text/image contracts before semantic search is .ready. Vision
remains a valid similarity service but can never satisfy semantic text search.
Qwen uses an internal QwenModelStatus with not-configured, checking,
available, missing, and invalid states. Settings translates it to the common
capability presentation and updates the Qwen feature at the same time.
Resource Managers and Their Cache
RawCullAIModelResourceManager<Provider> is an actor because filesystem
inspection, cryptographic bundle validation, and provider initialization must
not run on the main actor or race with a managed-location change.
The resource cache has two keys:
RawCullAIModelResourceSnapshot, a sorted list of path, file kind, byte
count, modification time, and symlink target; and- the resulting capability, optional provider, and optional provider-init
failure.
setManagedCandidateURL invalidates both when the URL changes. load() also
detects modifications within the same directory. The snapshot only decides
whether validation may be reused. PhotoAIKit’s resolver remains responsible for
metadata, required files, asset extension, fingerprint, and checksum validity.
Bundle validity and provider construction are reported separately. A bundle can
be structurally valid yet fail to initialize its concrete runtime; RawCull maps
that case to .invalid with the provider error so Settings can explain the
actual stage that failed.
Selecting the Similarity Runtime
RawCullAIModelRuntime.similarityService(prefersCLIP:clipModel:) has a strict
selection order:
- If the user disabled CLIP, return the existing Vision service.
- If the selected CLIP provider is absent, log the expected/resolved path and
return Vision.
- If the provider exists but its resolved location is missing, return Vision.
- Otherwise return a new
RawCullCLIPSimilarityService around the validated
provider and supply a factory that can reconstruct a provider from the exact
validated location for finite-vector recovery.
The service value can change while RawCullSimilarityFeature and
SimilarityScoringModel retain their identities. Backend descriptors determine
whether existing artifacts are still compatible.
Semantic search is constructed only from a currently validated CLIP provider.
The provider itself satisfies both TextEmbeddingProviding and
ImageTextSimilarityComparing, so RawCullCLIPSemanticSearchService can use
the same model identity as the cached image embeddings.
Installing and Replacing the Segmentation Runtime
The model runtime retains the selected RawCullSegmentationModel, currently
SAM 3. Selection and availability changes converge on
activateSelectedSegmentationProvider(availability:).
When a provider is available, installSegmentationProviderIfNeeded compares
its ModelIdentity with the active identity. A change rebuilds:
SubjectMaskRepositoryConfiguration with prompt, model identity, and maximum
input side;SubjectMaskRepository over the retained stores;SegmentationService over the new provider and stores; andSubjectMaskSelector over the matching repository and service.
When unavailable, the runtime installs the placeholder provider only if a real
identity was previously active. Identity checks prevent needless reconstruction
on repeated equivalent refreshes.
The stable DeepAIReviewFeature then receives:
- a new
RawCullDeepAIReviewPipeline when the provider exists; - a disk-mask loader when disk storage exists; and
- the current availability state.
If availability disappears during a review, DeepAIReviewFeature.install
cancels the active operation. Stored feature identity and already completed
results remain under one owner.
Objects Runtime: shared models, separate workflow
Objects uses the same validated SAM 3 provider as Deep Review, but calls its
instance-segmentation contract through a separate ObjectSegmentationService.
It uses the same QwenInferenceRuntime actor as standalone Qwen, but owns a
separate feature state machine and result array. These shared actors prevent
duplicate heavy model instances; the separate workflows keep subject-mask
selection and per-object instance analysis from sharing incompatible caches.
Activation and deactivation
The model runtime creates ObjectMaskMemoryStore at launch and tries
ObjectMaskDiskStore using the configured object-mask cache directory. It starts
with no object segmentation service. A complete managed-location snapshot
clears that service, changes SAM 3 and CLIP resource-manager candidates, and
validates or clears Qwen. Settings first cancels and marks Objects unavailable
while this snapshot is applied. The generation-gated capability refresh then
loads the SAM 3 provider; when available, it builds ObjectSegmentationService
with the object stores, 4,320-pixel maximum side, and eight-instance maximum.
Settings installs that service and the latest Qwen status into the existing
RawCullObjectAnalysisFeature.
The feature is ready only when both services exist. It distinguishes checking,
SAM 3 unavailable, Qwen unavailable, and both unavailable so the view can
direct users to the missing download. install(segmentation:qwenStatus:) cancels
an active batch when either dependency changes. Model removal therefore cannot
leave an active Objects operation attached to a stale provider.
Feature lifetime and task boundaries
RawCullApplicationState.make creates one object feature and passes it to
Settings, RawCullIntelligenceRuntime, and the AI Analysis view. Identity
assertions check that it shares the model runtime’s Qwen actor. The feature
retains the Qwen actor, image loader, mask stores, current object service,
availability, results, progress, and a cancellable task. It snapshots the
chosen concept mode, manual phrases, and criteria at batch start. Files are
processed sequentially, each concept is segmented sequentially, and
cancellation is checked between image loading, discovery, segmentation,
board construction, and Qwen assessment.
The model runtime is main-actor isolated for installation and selection.
ObjectSegmentationService and QwenInferenceRuntime are actors. The mask
deduplication and board rendering CPU work runs in concurrent tasks before
returning immutable results to the main-actor feature. Object results and
captured timings stay in memory; object masks may outlive a feature result in
the optional disk store. A detail view retrieves masks using the stored SAM 3
model identity and the same source/concept/cache parameters, then generates
yellow outlines without regenerating a model result.
Board rendering must produce a crop and matching mask crop for every retained
object. If either crop cannot be prepared, ObjectReviewBoardRenderer throws
reviewBoardUnavailable; the feature records a per-photo failure before Qwen
is asked to interpret the board. This keeps the board’s numbered panels and
the set of IDs requested from Qwen in agreement.
No-match, response failure, and retry
An empty retained instance set completes as No Matching Objects and skips the
Qwen board. When SAM 3 found instances but Qwen returned invalid or incomplete
assessment JSON, the result keeps the instances and free-form response, records
a stage-specific error, and remains eligible for Analyze or Retry Failed. It
does not become Complete merely because segmentation succeeded.
For assessment retry, the feature checks source size and modification date,
discovery mode, manual concepts, Qwen model name, and the stored result. It
loads every retained mask using PhotoAIKit’s object cache key, which also
encodes source identity, concept, SAM 3 model identity, input maximum side,
and instance limit. Only a complete cache hit reuses segmentation; otherwise
the workflow rediscovers concepts when needed and reruns SAM 3. Cancellation
and generation checks keep a superseded result from publishing.
The results table labels whole-photo Qwen confidence. The detail view labels
each SAM 3 mask score independently. The exact board-ID check establishes that
the response describes the expected number of objects; it cannot prove that
the descriptions correctly match those objects. The displayed object count is
the number of retained SAM 3 matches for the requested concepts, not a census
of the whole photograph. The September 24 puffin evaluation produced 16/16
structured results yet still exposed a swapped two-bird description and false
scene claims. In later in-app checks, _DSC3028.ARW showed opposing crop and
description associations for two birds, with the source of the mismatch still
unresolved. _DSC3031.ARW retained two birds but its Complete, 95%-confidence
Qwen summary invented a third. The detail panel displays only the currently
selected object’s assessment. See
AI Models in RawCull
for the prompt, mask filtering, board layout, timings, in-app observations, and
remaining validation work.
Qwen Runtime Lifetime
Qwen intentionally does not use the generic resource-manager actor. Its
QwenInferenceRuntime owns two lazy layers:
CoreAIQwenProvider, created during validation; andCoreAIVisionLanguageModel, created on the first assessment and retained for
subsequent sessions.
Every validation or clear increments modelGeneration. assess captures that
generation before an asynchronous model load and checks it after every
suspension. If Settings removes or replaces the model during loading or
generation, the old operation throws cancellation instead of publishing through
an obsolete model.
RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature each own a
separate batch generation and task while sharing that one inference actor.
Both process images one at a time, isolate per-file failures, and cancel when
their required model status changes. The features and inference actor protect
different races: batch/UI lifetime and provider/model lifetime.
Revisioned Configuration Application
Settings publishes one complete RawCullIntelligenceConfiguration containing:
- a monotonically increasing revision;
- the selected similarity service;
- semantic-search capability and optional service; and
- the selected segmentation model.
Its identity contains values, not provider references:
- selected similarity backend descriptor;
- accepted artifact backend descriptors;
- semantic capability;
- semantic backend descriptor; and
- segmentation selection.
Concrete services stay on @MainActor; only descriptor-based identity is
Sendable.
The apply algorithm
RawCullIntelligenceRuntime.apply(configuration:) follows this order:
- Compute incoming identity.
- Reject a revision less than or equal to the last accepted revision. A
same-revision/different-identity assertion catches a broken publisher.
- If a newer revision describes the same identity, record the newer revision
without resetting any work.
- If segmentation selection changed, ask the model runtime to activate it.
- If similarity backend or accepted artifact descriptors changed, replace the
similarity service through the stable feature.
- If semantic capability or semantic backend changed, replace semantic
configuration through that same stable similarity feature.
- Record the accepted identity and revision.
- Return the model runtime’s current capability snapshot to Settings.
sequenceDiagram
participant UI as Settings UI
participant Settings as RawCullAISettingsModel
participant Runtime as RawCullIntelligenceRuntime
participant Models as RawCullAIModelRuntime
participant Feature as RawCullSimilarityFeature
UI->>Settings: change CLIP/model/segmenter preference
Settings->>Settings: persist and increment revision
Settings->>Runtime: apply(complete snapshot)
Runtime->>Runtime: reject stale revisions and compare identity
opt segmentation changed
Runtime->>Models: setSelectedSegmentationModel
end
opt similarity changed
Runtime->>Feature: replaceSimilarityService
end
opt semantic configuration changed
Runtime->>Feature: replaceSemanticSearchConfiguration
end
Runtime-->>Settings: current capabilitiesRevision and identity solve different problems. Revision orders decisions;
identity determines whether a newer decision requires work.
What Service Replacement Invalidates
RawCullSimilarityFeature.replaceSimilarityService first compares complete
backend identities. For a real change it:
- asks the application context to cancel and reset burst analysis tied to the
old backend;
- installs the new service in
SimilarityScoringModel; - cancels existing image hydration;
- advances the image-hydration generation; and
- rehydrates the current catalog for the new accepted descriptors.
Semantic replacement updates semantic capability/service, cancels semantic
hydration, advances its independent generation, and rehydrates compatible
artifacts. Image similarity and semantic search use separate task handles and
generations so changing one concern does not confuse completion from the other.
Catalog hydration performs image and semantic hydration in order and finally
checks the current catalog identity. Ranking captures operation generation,
catalog identity, and backend identity. A late completion must match all three
before it is accepted.
Stable Identity: Why the Graph Is Not Rebuilt
Several views and owners retain the same feature objects:
RawCullApp ───────────→ RawCullIntelligenceRuntime
RawCullViewModel ─────→ RawCullSimilarityFeature A
Runtime ──────────────→ RawCullSimilarityFeature A
SwiftUI view ─────────→ RawCullSimilarityFeature A
On a model switch RawCull keeps A and changes its service. Rebuilding would
create a split graph where an existing view observes A while runtime commands
reach B. The objects might have the same type, but they would not share:
- active task handles;
- cancellation and generation state;
- indexing/search progress;
- hydrated artifacts and distances;
- Deep Review results and mask-candidate history;
- Qwen results and current batch;
- Objects results, numbered instance IDs, retry state, and mask-cache access;
- SwiftUI observation registrations; or
- application-context bindings.
Keeping the stateful owner stable also lets the owner make a precise
invalidation decision. Reconstructing everything would either lose unrelated
state or risk copying backend-specific state into an incompatible runtime.
Ownership and Weak Coordination Edges
The runtime strongly owns Settings, but Settings must call back to the runtime.
That callback is weak:
@ObservationIgnored private weak var configurationConsumer:
(any RawCullIntelligenceConfigurationApplying)?
AnyObject permits weak protocol storage. @ObservationIgnored prevents a
coordination detail from becoming UI state; it does not affect ownership.
Other coordination edges follow the same rule:
| Strong owner/child relation | Weak or per-run callback |
|---|
| Intelligence runtime → Settings | Settings → RawCullIntelligenceConfigurationApplying |
| Settings → Downloads model | Downloads model → RawCullAIManagedModelLocationsApplying |
| Runtime/view model → Similarity feature | Similarity feature → RawCullSimilarityApplicationContext |
| Runtime/view model → Semantic feature | Semantic feature → RawCullSemanticSearchApplicationTarget |
| Runtime/view model → Deep Review controller | Controller → DeepAIReviewApplicationContext |
| View model → Burst coordinator | Per-run closures capture [weak self] |
This produces one clear ownership direction and prevents retain cycles in both
the full application session and shorter-lived tests.
Actor Isolation and Work Placement
| Component | Isolation | Why |
|---|
RawCullAIModelRuntime | @MainActor | Atomically publishes capabilities and installs services used by observable features. |
RawCullIntelligenceRuntime | @MainActor | Applies ordered settings decisions to stable UI-facing objects. |
| Settings/features/controllers/scoring model | @MainActor | Own observable state, tasks, progress, and presentation. |
RawCullAIModelResourceManager | actor | Serializes location/cache state while keeping filesystem and provider setup off the main actor. |
CoreAICLIPProvider | actor | Owns lazy Core AI model/tokenizer state and serial inference. |
CoreAISAM3Provider | actor | Owns lazy segmentation engine/tokenizer state. |
SegmentationService | actor | Coordinates provider access and mask stores. |
ObjectSegmentationService | actor | Coordinates per-concept SAM 3 instance inference and the separate object-mask stores. |
QwenInferenceRuntime | actor | Owns provider, loaded VLM, and model generation. |
RawCullObjectAnalysisFeature | @MainActor | Owns availability, sequential batch, progress, results, retries, and generation. |
| Pure scoring/ranking functions | @concurrent or nonisolated | Run CPU-heavy work without making observable state unsafe. |
The main actor coordinates; it does not perform model hashing, model execution,
image vector comparison, or pixel-level subject-detail scoring itself.
Cancellation and Stale-Result Defences
RawCull uses several independent tokens because they protect different scopes:
| Counter or identity | Rejects |
|---|
Settings refreshGeneration | An older model/evidence refresh finishing after a newer one. |
| Runtime configuration revision | An older settings decision arriving after a newer decision. |
| Similarity hydration generations | Results from tasks invalidated by service or catalog changes. |
| Similarity ranking generation + catalog/backend identity | Ranking for an old anchor, catalog, or backend. |
| Deep Review generation | Progress/results after cancellation or restart. |
| Qwen feature generation | Batch results after cancellation/restart. |
| Objects feature generation | Object results after cancellation, model replacement, tool/source switch, or restart. |
| Qwen model generation | A lazy load or response using a removed/replaced provider. |
Task cancellation is cooperative, so the generation and identity checks are
essential. Cancellation requests work to stop; generations prevent late work
that did not stop immediately from becoming current state.
Failure and Fallback Policy
Runtime fallback is explicit:
- If CLIP is disabled, missing, invalid, or fails provider construction,
similarity uses Vision.
- Once a CLIP indexing pass begins, individual failures do not receive
Vision artifacts. Valid CLIP results remain; failed files stay unavailable.
- Semantic search is unavailable without a compatible text-capable CLIP
provider and compatible cached image artifacts.
- Deep Review is unavailable without the selected segmentation provider. A
failed candidate does not prevent later candidates from being evaluated.
- Disk mask-cache creation failure leaves memory storage usable, while the
capability reports the disk failure.
- Qwen absence clears its runtime; an invalid or text-only bundle is reported,
and an active batch is cancelled when status becomes unavailable.
- Objects requires both SAM 3 and Qwen. Its availability names the missing
dependency; it does not silently substitute CLIP, Vision, or generic prose.
An empty SAM 3 object set is a successful no-match result. An invalid Qwen
assessment after segmentation stays visible and retryable. A missing board
crop is reported as a per-photo failure before Qwen assessment.
This distinction between service-selection fallback and within-operation
fallback prevents heterogeneous artifacts and misleading results.
Runtime Lifecycle Summary
stateDiagram-v2
[*] --> GraphBuilt: construct stable graph
GraphBuilt --> VisionReady: synchronous initial configuration
VisionReady --> Checking: refresh installed locations
Checking --> ProvidersReady: validate bundles and construct providers
Checking --> PartialAvailability: some bundles missing or invalid
ProvidersReady --> Configured: publish newer configuration
PartialAvailability --> Configured: publish explicit capabilities/fallbacks
Configured --> Running: feature operations
Running --> Rechecking: download/remove/setting change
Rechecking --> Configured: identity-diffed apply
Configured --> [*]: app releases both stable rootsThe important invariant is that provider availability may change many times,
while application feature identities remain stable for the session.
Source Map
The runtime’s central rule is simple: validate and replace model-dependent
services behind stable state owners, then accept results only when revision,
generation, catalog, and backend identities still match.
11 - RawCull Packages
Pinned package revisions, imported products, dependency direction, and the recommended architecture reading order.
RawCull Packages
RawCull is the composition root for four architecture packages and four small
rsync/persistence support packages. The package repositories are separately
versioned. They are not copied source snapshots inside TechDocRawCull; paths
under Sources/ and Tests/ in these guides refer to the named package
repository. The revision notes at the start of each guide say whether its detailed
walkthrough was reviewed at the app’s current pin. In particular, the PhotoAIKit
and RawParserKit guides preserve their earlier architecture audits while the
lockfile table below records what the current app builds.
Resolved Dependency Snapshot
This table is derived from
RawCull.xcodeproj/project.xcworkspace/xcshareddata/swiftpm/Package.resolved.
This snapshot is from the current RawCull checkout on September 30, 2026.
“App” means RawCull links a product directly; “transitive” means another package
owns the dependency. A revision-only pin has no semantic-version label.
| Identity | Relationship | Version/branch | Revision |
|---|
| PhotoAIKit | app | revision pin | 77cc1d84a5d98a485caa15be102c8a55eb3d7698 |
| PhotoAnalysisKit | app | 1.3.1 | 2a1466e04d821fa2628d6985296643e0d0c7e465 |
| RawCullCore | app | 1.1.2 | d25a51e65ad32a82bf82f86fa0ec07d6e14498e9 |
| RawParserKit | app | 1.3.1 | f0e5b02a10294798afd86781da5d6510e146a7ba |
| RsyncArguments | app | 1.0.0 | 0ff6518136c208dfbecc1a918f045048ca79853d |
| RsyncProcessStreaming | app | 1.0.0 | dd86f012b352888fd146e0b6e103740dc237f740 |
| ParseRsyncOutput | app | 1.0.0 | e079e0c9d34bf07f7f2a4312b40feea79ae14847 |
| DecodeEncodeGeneric | app | 1.0.0 | b5ecbbbe1b244191efec1532a979f6ae342d6617 |
| coreai-models | transitive from PhotoAIKit | revision pin | 475c585fdb0fe82a83c8f777f259e9414bd44c98 |
| EventSource | transitive | 1.5.1 | 86b5096ac59ab46e66bd1f6377c604bc1dab0bc2 |
| swift-asn1 | transitive | 1.7.3 | 3b6410f7dee09eb33cdd26260c5fd47fda19b0e2 |
| swift-collections | transitive | 1.7.1 | 98ef3c98609a1e31b7e157b5b619579001a789d6 |
| swift-crypto | transitive | 4.5.2 | da9d28d69ebe3894b18376c8f2395c2f37b8448f |
| swift-huggingface | transitive | 0.11.0 | f2f99991f2d7d8fdb3187e4fd539cd2facf5c13d |
| swift-jinja | transitive | 2.5.1 | 4588064a20f3fc093c95f2f7d3359999bf30cae5 |
| swift-transformers | transitive | 1.3.4 | c21fdcde390313a6d98d8e33a346f2c3486c3ab0 |
| xgrammar | transitive | 0.2.2 | 4d145cc13d878c751ebeed36af1c013074be76bc |
| yyjson | transitive | 0.12.0 | 8b4a38dc994a110abaec8a400615567bd996105f |
Do not infer the app boundary from every transitive pin. Xcode product
references and source imports define what RawCull actually consumes.
Imported Products And Boundary Types
| Package | Products imported by RawCull | Values or protocols crossing the boundary |
|---|
| RawParserKit | RawParserKit | RawImageLoader, RawImageMetadata, RawFocusPoint, RawFormatRegistry, CGImage results; the app maps them through RawParserKitImageLoader |
| PhotoAnalysisKit | PhotoAnalysisKit | PhotoAnalyzer, PhotoAnalysisInput, PhotoAnalysisDescriptor, PhotoAnalysisResult, focus evidence and mask values |
| RawCullCore | RawCullCore | RawCullFileItem, RawCullSourceCatalog, ExifMetadata, burst inputs/results/configurations, ranking evidence, histograms; the app exposes compatibility typealiases such as FileItem |
| PhotoAIKit | PhotoAIContracts, CoreAICLIPBackend, CoreAISAM3Backend, CoreAIEfficientSAMBackend, VisionFeaturePrintBackend, PhotoAIWorkflows, PhotoAIStorage | AIImageSource, model identities/resources, similarity artifacts/descriptors, provider protocols, segmentation requests/results, mask stores and workflows |
| RsyncArguments | RsyncArguments | rsync argument builders used by Params and ArgumentsSynchronize |
| RsyncProcessStreaming | RsyncProcessStreaming | RsyncProcess and ProcessHandlers used by the copy executor |
| ParseRsyncOutput | ParseRsyncOutput | parsed transfer progress and totals used by RemoteDataNumbers |
| DecodeEncodeGeneric | DecodeEncodeGeneric | generic JSON encode/decode used by saved-file persistence |
RawCull owns the translations between these vocabularies. None of the four
architecture packages imports another merely to share an app model.
Dependency Direction
flowchart TD
UI["RawCull SwiftUI"] --> Host["RawCull composition, adapters, policy, persistence"]
Host --> Parser["RawParserKit\nRAW decode + normalized metadata"]
Host --> Analysis["PhotoAnalysisKit\nmeasurements + masks"]
Host --> AI["PhotoAIKit products\nAI contracts + backends + workflows"]
Host --> Core["RawCullCore\npure grouping + ranking"]
Host --> Args["RsyncArguments"]
Host --> Process["RsyncProcessStreaming"]
Host --> Output["ParseRsyncOutput"]
Host --> JSON["DecodeEncodeGeneric"]
Parser -. "app adapter" .-> Core
Analysis -. "app adapter" .-> Core
AI -. "app adapter" .-> Core
Args --> Process
Process --> OutputSolid arrows are compile-time imports by the app or support flow. Dotted arrows
are value translation performed by RawCull, not package dependencies.
Recommended Reading Order
- RawParserKit — a file becomes an orientation-normalized
image and neutral metadata.
- PhotoAnalysisKit — a decoded image becomes sharpness,
saliency, focus evidence, and masks.
- RawCullCore — measurements become burst boundaries,
recommendations, and review state.
- PhotoAIKit — optional CLIP similarity/semantic search and
segmentation run behind typed contracts.
- Return to the app pages to see RawCull compose those independent boundaries,
persist results, and present policy.
When a package pin changes, compare its manifest and public source at the new
resolved revision before updating this documentation. A sibling checkout may be
ahead of the revision RawCull actually builds.
11.1 - How PhotoAIKit Is Constructed
A detailed guide to PhotoAIKit’s contracts, CLIP image and text inference, semantic comparison, SAM 3 and EfficientSAM, workflows, storage, concurrency, and model identity.
How PhotoAIKit Is Constructed
Revision scope: This architecture walkthrough was written against
PhotoAIKit 20e57359603313af7c2d38cae3e8b6e37f8838ef. The current RawCull
checkout resolves 77cc1d84a5d98a485caa15be102c8a55eb3d7698.
Read the package behavior here as an architectural baseline and use
AI Models in RawCull and
The RawCull AI Runtime for the current app integration.
PhotoAIKit is a reusable Swift package extracted from application code. Its most
important achievement is not merely that CLIP, SAM 3, and EfficientSAM run. It
is that reusable AI behavior has been separated from RawCull’s UI, RAW-file
handling, paths, sandbox rules, and culling policy.
This document explains the construction from the bottom up and gives the reason
for each boundary.
1. Begin With The Package Boundary
The package owns:
- typed,
Sendable contracts; - model-bundle validation and fingerprinted model identity;
- Core AI CLIP image and text inference plus SAM 3 and EfficientSAM inference;
- validated image/text semantic comparison;
- Apple Vision feature-print generation and comparison;
- bounded similarity indexing and explicit fallback;
- segmentation, mask selection, geometry, and catalog workflows;
- optional artifact codecs and mask stores.
The host application owns:
- model download, installation, candidate URLs, and sandbox bookmarks;
- RAW decoding and application image models;
- SwiftUI, Observation state, and display wording;
- app-specific task staleness such as “latest selection wins”;
- query admission, result ordering, filtering, and semantic-search presentation;
- culling, burst ranking, sharpness, saliency, and rating policy;
- helper-process launch and application restart behavior.
This boundary makes PhotoAIKit reusable. A package that imports FileItem or
searches Bundle.main for a RawCull resource might be convenient for one app,
but it would silently encode application policy into the AI layer.
2. Read Package.swift As An Architecture Diagram
PhotoAIKit/Package.swift declares Swift tools 6.4, macOS 27, Swift 6 language
mode, seven library products, and one test target. It pins
apple/coreai-models to revision
cc812078731871574c9b2eb620aa40734c4b89ee in the audited revision;
RawCull now resolves coreai-models at
475c585fdb0fe82a83c8f777f259e9414bd44c98. The manifest declares
huggingface/swift-transformers from 1.3.3 (RawCull currently resolves 1.3.4).
flowchart TD
Contracts["PhotoAIContracts\nvalues + protocols"]
CLIP["CoreAICLIPBackend"] --> Contracts
Efficient["CoreAIEfficientSAMBackend"] --> Contracts
SAM3["CoreAISAM3Backend"] --> Contracts
Vision["VisionFeaturePrintBackend"] --> Contracts
Workflows["PhotoAIWorkflows"] --> Contracts
Storage["PhotoAIStorage"] --> Contracts
CoreAI["apple/coreai-models\nCoreAISegmentation product"] --> CLIP
CoreAI --> Efficient
CoreAI --> SAM3
Transformers["swift-transformers\nTokenizers product"] --> CLIP
Tests["PhotoAIKitTests"] --> Contracts
Tests --> CLIP
Tests --> Efficient
Tests --> SAM3
Tests --> Vision
Tests --> Workflows
Tests --> StorageThere is intentionally no large umbrella target in which every feature can reach
every implementation. Each higher-level product depends on PhotoAIContracts,
but the backends, workflows, and storage products do not depend on one another.
Why this shape helps:
- A host can import only the products it uses.
- Workflow code is testable with fake providers and decoders.
- Storage does not need to know which inference backend produced a value.
- A backend cannot accidentally reach into RawCull or into another backend.
- Framework-heavy dependencies remain concentrated in concrete backend targets.
3. PhotoAIContracts: The Stable Center
Contracts are the innermost layer. They contain data and protocol definitions,
not application decisions.
3.1 Translate Host Photos Into AIImageSource
Sources/PhotoAIContracts/AIImageSource.swift defines the package-owned source
value:
public struct AIImageSource: Codable, Hashable, Identifiable, Sendable {
public let id: UUID
public let url: URL
public let displayName: String
}
RawCull maps FileItem into this type at its integration boundary. PhotoAIKit
therefore receives only what a reusable indexing or segmentation workflow needs.
It never learns ratings, focus data, camera metadata, or view state.
SourceFileIdentity reads size and modification date. SourceFingerprint adds
the standardized path. These values let persisted artifacts answer a crucial
cache question: “Was this result produced for the current contents of this
source file?”
3.2 Invert Image Decoding
PhotoAIKit declares:
public protocol ImageDecoding: Sendable {
func image(for source: AIImageSource) async throws -> CGImage
}
The package needs a CGImage, but it should not prescribe how one is obtained.
A JPEG host can use ImageIO; RawCull can try RawParserKit first; tests can
return a generated image. The workflow depends on the capability, not on a
camera-format implementation.
This is the dependency-inversion principle in a small, concrete form:
flowchart LR
Workflow["PhotoAIKit indexer"] --> Protocol["ImageDecoding protocol"]
Raw["RawCull RAW decoder"] -. conforms .-> Protocol
Test["Test decoder"] -. conforms .-> Protocol3.3 Separate Generation From Comparison
Sources/PhotoAIContracts/SimilarityArtifact.swift defines two protocols:
ImageSimilarityArtifactProviding creates an artifact from a CGImage and
source.ImageSimilarityArtifactComparing computes the distance between two
compatible artifacts.
The combined ImageSimilarityBackend type alias requires both.
The split matters because generation can be actor-isolated and expensive, while
comparison may be synchronous and nonisolated. It also keeps distance semantics
with the backend that understands its payload. RawCull does not decode a
VNFeaturePrintObservation, nor does it implement CLIP cosine distance from
unverified arbitrary data.
3.4 Keep Text Queries Query-Scoped
Sources/PhotoAIContracts/TextEmbedding.swift defines a second, deliberately
different similarity value. TextEmbedding contains a normalized text vector
and a TextEmbeddingDescriptor with the complete backend identity, dimensions,
tokenizer version, and schema version.
Text embeddings are not file-backed SimilarityArtifact values. They have no
source fingerprint and are intended to live for one query. The split public
protocols mirror the image API:
TextEmbeddingProviding tokenizes and encodes a query.ImageTextSimilarityComparing compares a compatible image artifact with the
text embedding.ImageTextSimilarityBackend combines both capabilities.
The comparison returns cosine similarity in -1...1, where a larger value is a
closer semantic match. It is a relative retrieval score, not a confidence or a
keep/reject decision.
3.5 Make Persisted Artifacts Self-Describing
A SimilarityArtifact has two fields:
SimilarityArtifact
├── descriptor
│ ├── backend
│ ├── model fingerprint
│ ├── dimensions
│ ├── representation
│ ├── preprocessing version
│ ├── normalization version
│ ├── configuration version
│ ├── source fingerprint
│ └── schema version
└── payload (backend-owned Data)
This looks more elaborate than storing [Float], but it prevents subtle cache
bugs. A vector is not reusable merely because its dimension matches. Reuse is
valid only if the source, model asset, preprocessing, normalization,
configuration, representation, and schema are still the same.
SimilarityArtifactDescriptor.isCompatibleForDistance(with:) intentionally
compares every backend/configuration field but excludes the source fingerprint.
Two different photos must have different source fingerprints, yet their
artifacts can still be compared when they were produced by the same backend
definition.
3.6 Treat Model Identity As Data
The model types are spread across:
ModelIdentity.swift;ModelAssetFingerprint.swift;ModelBundleResolver.swift;ModelResource.swift.
ModelIdentity.cacheIdentifier preserves compatibility with older cache naming.
artifactIdentifier adds the selected asset fingerprint for new artifacts. This
distinction allows migration without pretending that two different model
binaries are the same model.
ModelResourceDescriptor.clip and .sam3 describe package-neutral requirements
and version strings. ModelCapabilityStatus reports available, missing, or
invalid without user-facing wording. The host translates that status into its
own settings UI.
4. Model Bundles Are Supplied, Not Discovered
PhotoAIKit never chooses an application directory. The host supplies one URL or
an ordered list of candidates.
A valid bundle has this conceptual layout:
ModelBundle/
├── metadata.json
├── tokenizer/
│ └── tokenizer.json
└── selected-model.aimodel (or .aimodelc)
metadata.json names the selected model in assets.main. New exports can also
provide asset_fingerprints.main.
ModelBundleResolver validates, in order:
- the URL exists and is a directory;
metadata.json decodes;assets.main is present and non-empty;- the asset extension is accepted;
- the selected asset exists;
- required resources such as
tokenizer/tokenizer.json exist; - the model asset can be fingerprinted;
- a manifest checksum, when present, matches the actual asset.
For a file asset, the cryptographic algorithm is SHA-256. For a compiled model
directory, PhotoAIKit hashes a stable sorted tree representation. Older
manifests without a checksum receive a size/modification-time fallback
fingerprint. That fallback can detect replacement in place but is marked as not
cryptographically verified.
ModelProviderFactory<Provider> connects validation to a backend constructor. A
backend supplies “how to make me from a validated URL”; the host supplies “which
URLs should be considered, and in what order.”
Candidate order is policy, not just presentation. ModelResourceResolver skips
a candidate that is missing, but returns immediately when it finds an invalid
candidate. This prevents a damaged higher-priority installation from being
silently hidden by a lower-priority fallback.
CLIP has a second validation layer after generic bundle resolution.
CLIPRuntimeConfiguration reads the model contract from metadata: source model
and revision, architecture, pretrained checkpoint, embedding dimensions,
preprocessing dimensions and normalization, tokenizer context and padding, named
image/text functions, normalization version, and configuration version. Current
bundles use shortest-side resize, a centered square crop, and bicubic
interpolation. Older PhotoAIKit bundles remain compatible through the original
224-pixel stretch and bilinear defaults.
This separation prevents a library update from unexpectedly changing an
application’s installation or sandbox policy.
5. Concrete Backend Products
5.1 CoreAICLIPBackend
CoreAICLIPProvider is an actor conforming to image embedding, artifact
generation/comparison, text embedding, and image/text comparison protocols.
Its responsibilities are deliberately backend-specific:
- validate the supplied CLIP bundle;
- lazily load and cache the named image and text Core AI functions;
- validate tensor names, shapes, scalar types, dimensions, and metadata;
- create the image, token, and attention-mask NDArrays;
- apply the model-declared resize/crop policy in sRGB and CLIP channel
normalization;
- L2-normalize image vectors through
ImageEmbedding; - validate that the exported graph’s text vector is finite, non-empty, correctly
shaped, and already L2-normalized;
- encode the vector as an artifact payload;
- compute cosine distance only after descriptor and payload validation.
The actor owns loadedModel, so lazy initialization is isolated from concurrent
callers.
For a semantic query, the same provider creates a TextEmbedding and
similarity(image:text:) rejects every incompatible model, representation,
dimension, preprocessing, normalization, configuration, tokenizer, schema, or
payload before computing a dot product. Vision artifacts are rejected because
they are not in CLIP’s image/text embedding space.
RawCull’s RawCullCLIPSemanticSearchService keeps the product policy outside
the backend: it trims and admits a literal query, selects already-persisted
compatible image artifacts, isolates per-file failures, reports progress, and
orders equal scores deterministically. It never decodes an image or persists a
query embedding.
5.2 CoreAIEfficientSAMBackend
CoreAIEfficientSAMProvider is an actor conforming to SubjectSegmenting. It
validates the EfficientSAM bundle, lazily owns ImageSegmenter, uses the
model’s empty point query (center point or regular grid depending on the
export), selects the highest-scoring mask, and adapts it to the same
SubjectSegmentationResult contract used by SAM 3. The pinned provider uses a
0.5 mask threshold and at most one segment.
5.3 CoreAISAM3Backend
CoreAISAM3Provider conforms to SubjectSegmenting. It owns tokenizer setup,
lazy model loading, request inference, query selection, confidence conversion,
mask decoding, resizing, thresholding, feathering, timing, and diagnostics.
The provider creates CoreAIClipTokenizer and asks Core AI for up to five
segments. Postprocessing prefers the exhaustive semantic probability map when
the exporter supplies one. Otherwise it unions every returned instance mask,
rather than selecting only the highest-scoring instance. The resulting
SubjectSegmentationResult therefore represents all subjects matching the
prompt. SAM3MultiSubjectTests protects both the semantic-map and instance-union
paths.
The public contract speaks in SubjectSegmentationRequest and
SubjectSegmentationResult, not Core AI tensors. This lets workflows and hosts
work at the domain level while the backend handles framework details.
5.4 VisionFeaturePrintBackend
This actor uses VNGenerateImageFeaturePrintRequest. The produced
VNFeaturePrintObservation is securely archived into an opaque artifact
payload.
Comparison returns to this backend, which unarchives the observations and calls
Vision’s native computeDistance. The observation never becomes part of
RawCull’s persistence API. This is an example of information hiding: callers
can store and route the artifact without learning its private representation.
6. PhotoAIWorkflows: Reusable Orchestration
Backends operate on one image or request. Workflows coordinate many sources and
reusable policies.
6.1 Similarity Indexing
EmbeddingIndexer is the vector-specific API. SimilarityArtifactIndexer is
the more general API used when the fallback might have an opaque representation
such as a Vision feature print.
Both indexers accept:
- package-owned sources;
- an injected decoder;
- a primary provider;
- an optional fallback provider;
- an explicit fallback policy;
- a concurrency limit;
- an async progress callback.
They use a throwing task group but enqueue only concurrencyLimit children.
Each time one finishes, one new source is added. This is bounded concurrency: a
catalog with 20,000 files does not create 20,000 live tasks and decoded images.
The fallback policies are:
| Policy | Behavior | Appropriate when |
|---|
.none | Keep primary successes and failures | There is no compatible fallback |
.perItem | Retry only a failed item | Primary and fallback results can safely coexist |
.wholeBatch | If any primary item fails, rerun every requested source through fallback | A batch must use one comparable representation |
The indexer does not decide which policy is correct for a host. RawCull’s
current CLIP service selects .none: it keeps successful CLIP artifacts and
reports per-file failures. RawCull selects Vision before indexing when CLIP is
disabled or no validated provider exists. The generic .wholeBatch mechanism
remains available to other hosts, but is not the current RawCull CLIP path.
6.2 Segmentation Workflows
The SAM 3 side is split into focused units:
| Type | Responsibility |
|---|
SegmentationService | Resize/preprocess coordination, cache lookup, inference, persistence, partitioning, and prefetch |
SegmentationBatchPipeline | Batch generation, progress events, summary, and versioned JSON-lines transport |
SubjectMaskRepository | Cache-only access using explicit prompt/model/size configuration |
SubjectMaskSelector | Ordered prompt fallback, quality/confidence thresholds, and best-mask selection |
SubjectMaskCatalogIndex | Incremental cache inventory without invoking inference |
SubjectMaskGeometry | Alpha coverage, bounds, centroid, and freshness |
SubjectMaskQuality | Reusable geometry-based quality classification |
The package can decide which mask best satisfies a general selection strategy.
RawCull still decides which photo is a culling candidate or winner. Reusable
mask quality is not the same concern as product ranking policy.
7. PhotoAIStorage: Optional Persistence Mechanics
Storage depends only on contracts.
For similarity data:
SimilarityArtifactCodec wraps an artifact in a versioned envelope.EmbeddingCodec writes descriptor-complete CLIP artifacts.decodeMigrating reads current data and two version-1 formats, but legacy
values are accepted only against the real current model identity and are
marked for immediate rewrite.LegacyCLIPEmbeddingCodec remains a staged migration reader; its
identity-incomplete writer is deprecated.
For segmentation:
SubjectMaskMemoryStore is an injected actor-backed dictionary.SubjectMaskDiskStore stores PNG mask data plus JSON metadata, validates the
storage key, reports disk usage, supports pruning, and rejects stale or
corrupt entries.
The host injects the disk directory. PhotoAIKit owns how the record is encoded
and validated, while the application owns where records live and when they are
removed.
8. Concurrency And Isolation
PhotoAIKit targets Swift 6 and makes boundary values Sendable.
The main patterns are:
- Actors for mutable runtime state: providers lazily cache loaded models;
services and stores protect mutable state.
- Value types for transport: sources, identities, descriptors, progress,
failures, and configuration values cross tasks safely.
- Bounded task groups: indexers and prefetch workflows limit simultaneous
work.
- Cooperative cancellation: workflow boundaries and expensive stages call
Task.checkCancellation(). - Async progress callbacks: the package reports neutral progress values; a
host decides how and where to publish UI state.
Actor isolation does not remove the need for a host policy. RawCull still owns
generation tokens and catalog checks that prevent an older task from replacing
results for a newer selection.
9. RawCull Integration And Policy Boundary
RawCull imports the package products and assembles concrete providers in
RawCullAIModelRuntime. RawCullApplicationState binds those providers to
stable features owned by RawCullIntelligenceRuntime:
| Package contract or implementation | RawCull adapter/consumer | Policy that remains in RawCull |
|---|
CoreAICLIPProvider, ImageSimilarityArtifactProviding, and ImageSimilarityArtifactComparing | RawCullCLIPSimilarityService, SimilarityScoringModel, and RawCullSimilarityFeature | managed model locations, selected CLIP model, RAW decoding, concurrency 1, retry/replacement recovery, per-file persistence, burst thresholds, and subject-mismatch adjustment |
VisionFeaturePrintBackend | RawCullVisionSimilarityService | always-available startup service, concurrency 4, service selection, cache admission, and UI state |
TextEmbeddingProviding and ImageTextSimilarityComparing | RawCullCLIPSemanticSearchService and RawCullSemanticSearchFeature | query admission, progress, catalog/rating filters, deterministic ties, result count, selection/navigation binding, and ephemeral query lifetime |
SubjectSegmenting, SegmentationService, mask stores, repository, and selector | RawCullAIModelRuntime, DeepAIReviewFeature, and DeepAIReviewController | SAM 3 versus EfficientSAM selection, application-support paths, saved-evidence status, candidate admission, review presentation, group-signature validation, and culling decisions |
The package owns validation, descriptors, backend actors, mathematical
comparison, bounded generic workflows, and reusable codecs/stores. The app owns
filesystem/security scope, camera decoding, settings, capability wording,
generation tokens, catalog identity, cache policy, culling rules, and UI. In
particular, PhotoAIKit does not know FileItem, a burst winner, or where
RawCull installs a model. RawCullIntelligenceRuntime owns stable feature
lifetimes and applies revisioned configuration; views consume the focused
features and controller rather than low-level package providers.
10. Testing The Architecture, Not Only The Math
Most of Tests/PhotoAIKitTests/ exercises the products through public APIs. The
CLIP text suite also uses @testable import for deterministic tokenizer, batch,
tensor-shape, and preprocessing checks that sit below the provider boundary.
Together the tests cover architectural promises such as:
- model bundles are accepted through supplied URLs;
- manifest fingerprints become artifact identity;
- provider factories and resource resolvers share validation;
- CLIP normalization and cosine distance remain stable;
- model-declared CLIP preprocessing and legacy preprocessing remain compatible;
- token batches, attention masks, text-output validation, and image/text
compatibility checks reject malformed or mismatched data;
- the generic whole-batch fallback policy produces a homogeneous result set
(RawCull currently uses
.none for CLIP); - Vision artifacts stay opaque and use the native metric;
- segmentation caches by package-owned source values;
- disk stores reject stale or corrupt entries;
- legacy data is a rewrite candidate, not silently current data;
- batch transport has an explicit versioned schema;
- cancellation crosses the service boundary.
- SAM 3 preserves all matching subjects through its semantic-map or union fallback.
Fake decoders, providers, and stores make these tests possible. That testability
is a direct consequence of putting protocols in the innermost target.
The package includes tools under Tools/ for exporting CLIP and SAM 3 assets,
selecting a SAM 3 asset, generating fingerprints, and producing reference CLIP
image/text similarities. Output paths are explicit.
export_clip.py supports the existing OpenAI clip-vit-base-patch32 model and
OpenCLIP ViT-B-32-256 with datacomp_s34b_b86k weights. It writes the image
and text encoders as named functions in one Core AI asset, verifies tokenizer
parity for the OpenCLIP export, and records the runtime metadata consumed by
CLIPRuntimeConfiguration.
The Swift package itself declares no model resources. It does not embed
.aimodel, .aimodelc, tokenizer, or metadata assets. Model binaries are large
deployment inputs with their own lifecycle; keeping them out of the library
avoids coupling package source, application installation, and model
distribution.
12. How To Add Another Similarity Backend
Use the existing layering as a checklist:
- Add or reuse package-neutral contract values in
PhotoAIContracts; do not
add host types. - Create a separate backend target that depends on contracts.
- Conform to artifact providing and comparing protocols.
- Give the backend a complete
SimilarityBackendDescriptor. - Put framework-specific payload encoding and distance semantics inside that
backend.
- If the backend supports text, add a query-scoped descriptor and keep
image/text compatibility checks with that backend.
- Reuse
SimilarityArtifactIndexer and inject the host’s decoder. - Decide explicitly whether fallback is none, per-item, or whole-batch.
- Add public-API tests with fake sources and model-free inputs where possible.
- Let the host choose model URLs, cache locations, settings, and product
policy.
If a proposed type needs SwiftUI, FileItem, a RawCull path, or a burst rating,
it probably belongs in RawCull’s adapter layer instead.
Source Map
| Topic | PhotoAIKit source |
|---|
| Products and dependency graph | Package.swift |
| Host-neutral image and decoder boundary | Sources/PhotoAIContracts/AIImageSource.swift, ImageEmbedding.swift |
| Model validation and identity | Sources/PhotoAIContracts/ModelBundleResolver.swift, ModelResource.swift, ModelIdentity.swift, ModelAssetFingerprint.swift |
| Similarity descriptors and protocols | Sources/PhotoAIContracts/SimilarityArtifact.swift |
| Text embeddings and image/text comparison | Sources/PhotoAIContracts/TextEmbedding.swift |
| Segmentation contracts | Sources/PhotoAIContracts/SubjectSegmentation.swift, SubjectMaskStorage.swift |
| CLIP runtime and backend | Sources/CoreAICLIPBackend/CLIPRuntimeConfiguration.swift, CoreAICLIPProvider.swift |
| EfficientSAM backend | Sources/CoreAIEfficientSAMBackend/CoreAIEfficientSAMProvider.swift |
| SAM 3 backend | Sources/CoreAISAM3Backend/CoreAISAM3Provider.swift |
| Vision backend | Sources/VisionFeaturePrintBackend/VisionFeaturePrintBackend.swift |
| Similarity orchestration | Sources/PhotoAIWorkflows/EmbeddingIndexer.swift, SimilarityArtifactIndexer.swift |
| Segmentation orchestration | Sources/PhotoAIWorkflows/SegmentationService.swift, SegmentationBatchPipeline.swift, SubjectMask*.swift |
| Codecs and stores | Sources/PhotoAIStorage/ |
| Package boundary audit | Documentation/ExtractionMap.md |
| Public behavior tests | Tests/PhotoAIKitTests/ |
| RawCull semantic-search policy | RawCull/Intelligence/SemanticSearch/RawCullSemanticSearchService.swift, RawCull/Intelligence/SemanticSearch/RawCullSemanticSearchFeature.swift, RawCull/Intelligence/Similarity/SimilarityScoringModel.swift |
Next, follow these abstractions into the host application in
How RawCull Enables and Uses CLIP.
11.2 - How PhotoAnalysisKit Is Constructed
A detailed guide to PhotoAnalysisKit’s image-analysis boundary, sharpness pipeline, focus evidence, masks, calibration, batching, feature prints, resources, and concurrency.
How PhotoAnalysisKit Is Constructed
Revision audited: RawCull resolves PhotoAnalysisKit 1.3.1 at
2a1466e04d821fa2628d6985296643e0d0c7e465. The facade, descriptors, presets,
batch behavior, calibration, evidence, and mask APIs below describe that
revision.
PhotoAnalysisKit is the reusable measurement layer extracted from RawCull. It
turns a decoded CGImage plus neutral capture metadata into sharpness,
saliency, focus evidence, and an optional focus-mask image. It can also create
opaque Apple Vision feature prints.
The package measures a photo; it does not decide whether to keep it. That
distinction prevents image-processing code from becoming coupled to RawCull’s
files, settings, views, caches, or culling policy.
1. Begin With The Package Boundary
PhotoAnalysisKit owns:
- sRGB normalization for predictable analysis input;
- Core Image and Metal Laplacian processing;
- Vision saliency and optional subject classification;
- scalar sharpness metrics and failure classification;
- AF-point, subject, local-patch, and global focus evidence;
- focus-mask selection and rendering;
- numeric configuration, presets, quality levels, and analysis identity;
- bounded batch analysis and calibration;
- Vision feature-print generation and native comparison.
The host application owns:
- RAW, JPEG, or other source decoding;
- file URLs, security-scoped access, and sandbox bookmarks;
- source identity, cache keys, cache directories, and persistence;
- observable task state, progress presentation, and cancellation policy;
- settings labels and saved-settings migration;
- sorting, ratings, burst decisions, and culling actions.
The input boundary is intentionally small:
flowchart LR
Source["RAW or rendered source"] --> Decoder["Host decoder"]
Decoder --> Input["PhotoAnalysisInput\nCGImage + ISO + aperture + AF point"]
Input --> Analyzer["PhotoAnalyzer"]
Analyzer --> Result["PhotoAnalysisResult\nsaliency + breakdown + optional mask"]
Result --> Host["Host cache, UI, and culling policy"]PhotoAnalysisKit does not import RawParserKit or RawCullCore. RawCull provides
the adapters between them.
2. Read Package.swift As The First Design Document
The manifest declares Swift tools 6.2, Swift 6 language mode, macOS 26, one
library product, and one test target.
PhotoAnalysisKit package
├── PhotoAnalysisKit library
│ ├── public contracts and facade
│ ├── Core Image, Vision, and Accelerate implementation
│ └── packaged default.metallib
└── PhotoAnalysisKitTests
There are no third-party package dependencies. The implementation uses Apple
frameworks including Core Graphics, Core Image, Vision, Accelerate, and
Foundation.
The target enables InferIsolatedConformances and
NonisolatedNonsendingByDefault. The package is therefore built under the same
strict Swift 6 concurrency assumptions as the host without adopting UI
isolation.
Sources/PhotoAnalysisKit/PhotoAnalysisInput.swift defines the package’s source
value:
public struct PhotoAnalysisInput: Sendable {
public let image: CGImage
public let iso: Int
public let aperture: Double?
public let normalizedAFPoint: CGPoint?
}
The AF point uses normalized 0...1 coordinates with the origin at the visual
top-left. This matches the camera metadata shape used by RawCull. Vision uses a
bottom-left coordinate system, so the package performs the vertical conversion
internally.
ISO is clamped to at least 1. Aperture and AF point are optional because not
every image or camera provides them. The package can still compute a global
result when either is absent.
This value has no URL or file identifier. Two consequences follow:
- The package cannot reopen a source behind the host’s back.
- The host must associate returned results with the correct file and invalidate
them when that file or decode policy changes.
RawCull makes that ownership concrete in RawCullPhotoAnalysisAdapter. The
adapter chooses an embedded-preview decode through RawParserKit or a host-owned
RAW demosaic, bounds the requested pixel size, and then constructs
PhotoAnalysisInput. It also translates the returned neutral saliency value
into RawCullCore’s SaliencyInfo. Neither package needs to import the other to
participate in that flow.
PhotoAnalysisResult returns:
| Field | Meaning |
|---|
saliency | Optional neutral subject label and confidence |
breakdown | Scalar score plus detailed evidence and diagnostics |
focusMask | Optional rendered overlay image |
score | Convenience access to breakdown.finalScore |
An absent breakdown means analysis could not produce a valid result or was
cancelled. It is different from a valid score of zero.
4. PhotoAnalyzer Is The Public Facade
Sources/PhotoAnalysisKit/PhotoAnalyzer.swift keeps the public entry points
small:
| API | Work performed |
|---|
analyze | Saliency, classification, scalar scoring, and focus evidence; no overlay rendering |
analyzeWithFocusMask | Analysis plus a mask derived from the same evidence |
focusMask | Mask rendering, optionally reusing previously computed evidence |
calibrate | Catalog-sample calibration of the visual edge threshold |
sharpnessDescriptor | Stable identity for cacheable non-mask analysis behavior |
The facade contains an immutable FocusMaskEngine. Both types are
@unchecked Sendable because CIContext does not declare sendability even
though this package holds no mutable model or UI state and reuses the context
concurrently.
For each input, PhotoAnalyzer copies the supplied configuration, replaces its
ISO with the input ISO, and derives an aperture hint from the input aperture.
Callers can safely reuse one base configuration across many files.
When a host already stores SharpnessBreakdown.focusEvidence, passing it to
focusMask avoids repeating the saliency selection step. This makes the
measurement result useful as an input to later presentation work.
5. Follow The Scalar Sharpness Pipeline
The core implementation is in FocusMaskEngine+Scoring.swift.
flowchart TD
Image["CGImage"] --> SRGB["8-bit sRGB normalization"]
SRGB --> Vision["Vision saliency + optional classification"]
SRGB --> Preblur["ISO- and aperture-aware Gaussian pre-blur"]
Preblur --> Laplacian["Metal focusLaplacian kernel"]
Laplacian --> Samples["Global, saliency, AF, and local-patch samples"]
Vision --> Samples
Samples --> Tail["Robust p90-p97 tail scores"]
Tail --> Blend["Subject/global blend"]
Blend --> Adjust["Silhouette, subject-size, and blur-gate adjustments"]
Adjust --> Breakdown["SharpnessBreakdown"]5.1 Normalize Before Measuring
normalizeToSRGB redraws the input into an 8-bit sRGB RGBA bitmap. The Metal
pipeline therefore receives a predictable pixel format regardless of the source
image’s bit depth or color space.
5.2 Find Candidate Subject Regions
VNGenerateAttentionBasedSaliencyImageRequest supplies salient-object
rectangles. Small weak objects are removed, and candidates are ordered using AF
overlap or distance, saliency confidence, interior detail, area, and
deterministic coordinates.
When classification is enabled, VNClassifyImageRequest supplies a neutral
subject label. The package filters out broad environment descriptions and
prefers likely subjects such as animals or people. This label is evidence, not a
keep/reject decision.
5.3 Build Edge Energy
The packaged focusLaplacian Metal kernel runs after Gaussian pre-blur. The
pre-blur grows with ISO so high-frequency sensor noise is less likely to
masquerade as detail. Aperture hints damp or strengthen parts of this behavior
for wide, middle, and landscape apertures.
The result is rendered as floating-point RGBA data. The red channel carries edge
energy.
5.4 Score Several Regions
The engine samples:
- the full image after excluding a configurable border;
- the selected salient region;
- the camera AF region;
- smaller AF-center and AF-neighborhood regions;
- ranked local patches within subject or AF regions.
robustTailScore sorts the sample values, measures the p90-p97 energy band
relative to a p20 noise floor, and penalizes a band that is too sparse.
microContrast computes the standard deviation of finite edge samples.
The normal blend combines full-frame and subject evidence. When both AF and
saliency scores exist, AF has the larger share of the subject score. A
conservative local-detail component can then refine the broad subject
measurement.
The final value may be adjusted for:
- a silhouette-dominated subject whose strongest energy is mostly at the outer
rim;
- the area of a saliency-only subject;
- an aperture-aware soft blur gate driven by subject micro-contrast.
SharpnessBreakdown preserves the component scores and selected evidence so a
host can explain the result instead of presenting a single unexplained number.
5.5 Distinguish Failure Shapes
FocusFailureKind classifies the evidence as:
.motionBlur when global, subject, and AF detail are all weak and
micro-contrast is low;.missedFocus when the frame has usable global detail but the subject is much
weaker;.none when neither pattern is established.
These are algorithmic classifications, not final user-facing copy or culling
actions.
Scalar scoring answers “how much reliable detail is present?” Mask rendering
answers “where should the UI draw visible focused edges?” They share evidence
but have different configuration needs.
FocusEvidence records the winning region, AF scores, selected patch rankings,
spatial alignment, dominance, silhouette handling, visual threshold, coverage,
and confidence diagnostics. FocusPatchRanking exposes the detail, coverage,
shape, position, and penalty components behind each candidate patch.
The overlay pipeline:
- chooses AF-center, AF-neighborhood, AF, saliency, mixed, or global evidence;
- builds one native-pixel fine-detail Laplacian with clamped edges for every region;
- ranks local patches and selects the best evidence patches;
- chooses an adaptive percentile threshold from the complete selected search regions;
- applies optional erosion and dilation to the binary edge mask;
- colorizes, clips to the full search regions, feathers, and crops the mask;
- measures visible coverage after rendering with GPU reductions;
- returns updated evidence and render diagnostics with the image.
The renderer deliberately does not lower the threshold to force visible pixels.
guaranteeVisibleFocusEvidence remains in the public configuration for source
compatibility, but a weak or unfocused image may produce an empty mask.
FocusMaskRegionSource describes whether saliency, AF, both, or neither
provided the overlay region. FocusEvidenceOverlayStyle distinguishes subject
edges from global edges. RawCull decides how these neutral values appear in its
interface.
7. Configuration, Presets, And Cache Identity
SharpnessConfiguration is a value snapshot containing numeric algorithm
settings. It separates host-editable behavior from observable settings models.
Notable groups are:
| Configuration group | Examples |
|---|
| Edge pipeline | preBlurRadius, threshold, energyMultiplier |
| Mask morphology | dilationRadius, erosionRadius, featherRadius |
| Visibility | guaranteeVisibleFocusEvidence, minimumEvidenceCoverage |
| Region selection | AF radii, border inset, saliency weight |
| Scoring adjustments | subject-size factor, silhouette strength, fine-detail weight |
| Capture hints | iso, apertureHint |
SharpnessPreset applies high-level subject tuning for automatic, birds and
wildlife, portrait, landscape, or general action use. SharpnessQuality selects
fast, balanced, or high-precision fine-detail work. The host may persist its own
UI enums, but should map them into these package values instead of duplicating
constants.
Persisted results need more than a filename. SharpnessAnalysisDescriptor
records:
- descriptor schema version;
- scalar algorithm version;
- ISO-scaling policy version;
- aperture-hint policy version;
- every host-configurable value that affects non-mask scoring;
- the stable scoring energy multiplier.
The descriptor deliberately excludes per-image ISO and aperture, mask-only
presentation settings, decoded dimensions, source choice, and source-file
identity. A host must add those values to its cache identity.
Host sharpness cache identity
├── PhotoAnalysisKit SharpnessAnalysisDescriptor
├── per-image ISO and aperture
├── source file identity
├── selected preview/decode policy
└── decoded pixel dimensions
8. Bounded Batch Analysis
PhotoAnalysisBatchRequest<Identifier> contains an identifier and an async
input provider. The provider inversion is important: PhotoAnalysisKit
coordinates work, while the host retains file access and decoding policy.
analyzeBatch starts only up to maximumConcurrentTasks child tasks. When one
completes, it enqueues one more request. This prevents a large catalog from
creating an unbounded number of decoded bitmaps.
The API has three ordering and failure guarantees:
- progress is emitted in completion order;
- the returned results preserve request order;
- a decode failure is represented by a result whose
analysis is nil.
If the parent task is cancelled, the method cancels the group and returns nil,
telling the host to discard partial results rather than mistake them for a
complete batch.
9. Calibration Changes The Overlay, Not The Score
Calibration samples Laplacian energies from host-provided decoded inputs with
bounded concurrency. It downsamples the collected sample set, sorts it, and
returns p50, p90, p95, and p99 statistics plus a clamped threshold at the
requested percentile.
At least minimumSuccessfulImages inputs must succeed. Cancellation or too few
samples returns nil.
Calibration changes only the visual edge threshold. Core sharpness scores keep a
stable gain and do not depend on the current catalog. This prevents the same
photo from receiving a different scalar score merely because unrelated photos
were added or removed.
10. Vision Feature Prints Stay Opaque
VisionFeaturePrintBackend is an actor that creates a
VNFeaturePrintObservation using Vision revision 2 by default. The observation
is securely archived inside VisionFeaturePrint.payload.
The value also stores its Vision revision and representation version. Before
comparison, the backend verifies both values, securely unarchives each
observation, and calls Vision’s native computeDistance.
The host can persist the opaque payload, but it must associate it with
source-file identity. Incompatible prints return nil; corrupt archives or
failed Vision operations throw a typed VisionFeaturePrintError.
PhotoAnalysisKit intentionally does not copy PhotoAIKit’s general similarity
artifact descriptors, CLIP fallback, or batch indexing. It supplies only the
focused Vision measurement primitive.
Sources/PhotoAnalysisKit/Resources/Kernels.ci.metal is the source for the Core
Image kernel. SwiftPM copies Metal source resources but does not compile them
for command-line builds, so the package also checks in default.metallib.
Tools/build_metallib.sh regenerates the binary. A kernel change is incomplete
until the checked-in library has been rebuilt and tests have been run. The
binary resource affects algorithm output and must be treated like source, not an
optional deployment file.
12. Concurrency And Cancellation
The package uses four complementary techniques:
- Immutable facade and engine: callers pass input and configuration
snapshots; no app state is retained.
- Explicit concurrent workers: synchronous Core Image and Vision work runs
outside caller UI isolation while retaining task priority and task-local
context.
- Bounded task groups: batch analysis and calibration cap simultaneous
inputs.
- Cooperative checks: expensive stages test cancellation before Vision work,
rendering, large loops, and result publication.
@unchecked Sendable is confined to the immutable Core Image facade and engine.
Public transport values are Sendable value types.
13. RawCull Call Maps
The app has two deliberate entry paths into the same PhotoAnalyzer facade.
flowchart LR
Settings["Settings + photo-type preset"] --> Sharp["SharpnessScoringModel"]
Sharp --> Adapter["RawCullPhotoAnalysisAdapter"]
Adapter --> Input["PhotoAnalysisInput: decoded CGImage + neutral metadata"]
Input --> Batch["PhotoAnalyzer.analyzeBatch"]
Batch --> Results["PhotoAnalysisResult per FileItem"]
Results --> Core["RawCullCore ranking evidence + BurstAnalysisCache"]SharpnessScoringModel applies the selected package preset, builds requests
through RawCullPhotoAnalysisAdapter, and calls the facade’s bounded batch API.
The adapter owns source loading and file-to-identifier mapping; the package owns
measurement and its descriptor. BurstAnalysisCache records
PhotoAnalyzer.sharpnessDescriptor(for:), so a configuration-identity change
invalidates incompatible cached scores.
flowchart LR
UI["Focus overlay or calibration action"] --> Model["FocusMaskModel"]
Model --> Analyze["PhotoAnalyzer.analyzeWithFocusMask / focusMask / calibrate"]
Analyze --> Evidence["FocusEvidence + FocusCalibrationResult + CGImage mask"]
Evidence --> Presentation["RawCull FocusMaskResult + overlay state"]FocusMaskTypes.swift aliases neutral package types such as FocusEvidence,
FocusFailureKind, and FocusCalibrationResult, then layers RawCull
presentation metadata over SharpnessBreakdown. FocusMaskModel consumes the
returned mask image on app-owned isolation; scoring and evidence remain package
values.
14. Testing The Public Contract
The executable contracts are split by concern:
| Package test | Required behavior |
|---|
Tests/PhotoAnalysisKitTests/SharpnessMetricsTests.swift | Scalar metric and normalized sharpness calculations |
Tests/PhotoAnalysisKitTests/PhotoAnalysisBatchTests.swift | Bounded batch completion, ordering, failure isolation, and cancellation |
Tests/PhotoAnalysisKitTests/SharpnessAnalysisDescriptorTests.swift | Stable configuration identity and descriptor changes |
Tests/PhotoAnalysisKitTests/PhotoAnalyzerTests.swift | Facade behavior, focus evidence, calibration, and mask rendering |
Tests/PhotoAnalysisKitTests/VisionFeaturePrintTests.swift | Opaque Vision feature-print creation and comparison |
Tests/PhotoAnalysisKitTests/ imports the library through
import PhotoAnalysisKit, not @testable import. Tests therefore exercise the
public surface used by a host.
The suite covers:
- end-to-end sharpness analysis with the packaged Metal kernel;
- focus-mask images, evidence, and diagnostics;
- calibration from decoded images and neutral metadata;
- bounded concurrency, order preservation, progress, decode failure, and
cancellation;
- descriptor encoding and changes in identity-affecting settings;
- robust-tail, micro-contrast, ISO scaling, and focus-failure metrics;
- configuration presets and aperture behavior;
- Vision feature-print compatibility, round trips, and malformed payloads.
Synthetic images keep the suite deterministic and independent of RAW files,
model downloads, cache directories, and application state.
15. How To Add Another Analysis
Use the existing boundary as a checklist:
- Accept
CGImage and only the neutral metadata the algorithm needs. - Keep URLs, camera decoding, settings state, and source selection in the host.
- Return a
Sendable package-owned result with enough evidence to explain it. - Put numeric tuning in a value configuration, not UI preferences.
- Define a versioned descriptor when hosts may persist the output.
- Make synchronous framework work cancellation-aware and independent of UI
isolation.
- Bound multi-image work instead of spawning one live task per catalog item.
- Test through the public API with synthetic images and values.
If the proposed API needs FileItem, SwiftUI, a cache directory, or a rating,
that concern belongs in RawCull’s adapter or policy layer.
Source Map
| Topic | PhotoAnalysisKit source |
|---|
| Product and resource declaration | Package.swift |
| Input and result boundary | Sources/PhotoAnalysisKit/PhotoAnalysisInput.swift |
| Public analysis facade and metrics | Sources/PhotoAnalysisKit/PhotoAnalyzer.swift |
| Batch orchestration | Sources/PhotoAnalysisKit/PhotoAnalysisBatch.swift |
| Configuration and presets | Sources/PhotoAnalysisKit/SharpnessConfiguration.swift, SharpnessPresets.swift |
| Cache descriptor | Sources/PhotoAnalysisKit/SharpnessAnalysisDescriptor.swift |
| Public evidence values | Sources/PhotoAnalysisKit/FocusMaskTypes.swift |
| Engine isolation | Sources/PhotoAnalysisKit/FocusMaskEngine.swift |
| Saliency and scalar scoring | Sources/PhotoAnalysisKit/FocusMaskEngine+Scoring.swift |
| Overlay generation and patch ranking | Sources/PhotoAnalysisKit/FocusMaskEngine+MaskGeneration.swift |
| Calibration | Sources/PhotoAnalysisKit/FocusMaskCalibration.swift |
| Vision feature prints | Sources/PhotoAnalysisKit/VisionFeaturePrintBackend.swift |
| Metal source and compiled resource | Sources/PhotoAnalysisKit/Resources/ |
| Extraction decisions | Documentation/ExtractionMap.md |
| Public behavior tests | Tests/PhotoAnalysisKitTests/ |
| RawCull decode and result adapter | RawCull/Model/ViewModels/FocusandSharpness/RawCullPhotoAnalysisAdapter.swift |
Next, see how package-neutral measurements become culling-domain decisions in
How RawCullCore Is Constructed.
11.3 - How RawCullCore Is Constructed
A detailed guide to RawCullCore’s package-safe models, capture-time and EV-aware burst grouping, ranking evidence, confidence rules, histograms, concurrency, and tests.
How RawCullCore Is Constructed
Revision audited: RawCull resolves RawCullCore 1.1.2 at
d25a51e65ad32a82bf82f86fa0ec07d6e14498e9.
RawCullCore is the small domain layer at the center of RawCull. It does not
decode or analyze photos. Instead, it receives file metadata and measurements
already produced elsewhere, groups sequential files into bursts, and ranks the
candidates using deterministic culling rules.
Its main architectural value is that a recommendation can be tested without
opening a RAW file, running Vision, creating a Metal context, loading
application settings, or constructing a view model.
1. Begin With The Package Boundary
RawCullCore owns:
- package-safe file, EXIF, source-catalog, and saliency values;
- normalized focus-point parsing;
- burst group, boundary-evidence, ranking, confidence, and review-state models;
- pure burst grouping rules;
- pure burst ranking and one-click-safety rules;
- a lightweight 256-bin luminance histogram.
RawCullCore deliberately does not own:
- TIFF, MakerNote, ARW, or NEF parsing;
- thumbnail, preview, or full RAW decoding;
- Vision, Core Image, Metal, CLIP, or SAM 3 analysis;
- caches, settings, security-scoped URLs, and saved-file coordination;
- observable view models, SwiftUI, selection, sorting, and culling actions.
This produces a clear flow:
flowchart LR
Parser["RawParserKit\nmetadata + AF location"] --> Adapter["RawCull adapter"]
Analysis["PhotoAnalysisKit\nsharpness + saliency"] --> Adapter
AI["PhotoAIKit\nvisual distance"] --> Adapter
Adapter --> Values["RawCullCore values"]
Values --> Group["BurstGroupingEngine"]
Group --> Rank["BurstRankingEngine"]
Rank --> Decision["RawCull presentation and action"]The arrows describe how RawCull composes runtime values. RawCullCore imports
none of the other packages.
2. The Manifest Optimizes For A Small Domain Library
Package.swift declares Swift tools 6.2, Swift 6 language mode, macOS 26, one
RawCullCore library target, and one test target. It has no external package
dependencies and no bundled resources.
The target uses:
.defaultIsolation(MainActor.self)
.enableUpcomingFeature("InferIsolatedConformances")
.enableUpcomingFeature("NonisolatedNonsendingByDefault")
Default main-actor isolation matches the surrounding application environment,
but the package’s public value types and pure engines are explicitly
nonisolated. Background grouping and ranking therefore do not acquire
unnecessary main-actor hops.
This is a useful distinction: the build default is conservative, while each
proven pure API opts out explicitly.
3. Package-Safe Models Replace Application State
The package uses transport values instead of importing RawCull’s FileItem or
view models.
ExifMetadata is a Codable, Hashable, and Sendable snapshot containing
display strings and numeric values for shutter speed, focal length, aperture,
ISO, exposure compensation, camera, lens, RAW type, size class, and pixel
dimensions.
It contains both strings and numbers because they serve different purposes:
- strings preserve display-ready metadata such as shutter notation and lens
names;
- numeric exposure time, focal length, aperture, ISO, and exposure compensation
support deterministic stop-based rules without reparsing display text.
RawCull or an adapter constructs this value from RawParserKit’s
RawImageMetadata. RawCullCore does not know where the values came from.
At the app boundary, RawCull/Main/RawCullFileItem.swift deliberately preserves
familiar app names:
typealias FileItem = RawCullFileItem
typealias ARWSourceCatalog = RawCullSourceCatalog
typealias ExifMetadata = RawCullCore.ExifMetadata
These aliases are migration conveniences, not permission for the package to
import application state. Feature code may say FileItem, while the stored
value and package API remain RawCullFileItem.
3.2 RawCullFileItem
RawCullFileItem contains an ID, URL, name, byte size, modification date,
optional capture date and capture-time-zone offset, optional EXIF snapshot, and
optional normalized AF point.
Equality and hashing use only id. Two snapshots with the same ID are the same
logical file item even if another stored field changed. This matches the
identity behavior expected by RawCull collections.
effectiveCaptureDate prefers the EXIF capture instant and falls back to the
file modification date. usesFileModificationDateForCaptureTime exposes which
path was used so grouping and confidence can treat the fallback conservatively.
formattedSize is a lightweight Foundation convenience. File access and mutable
scan state remain outside the value.
3.3 Catalog And Saliency Summaries
RawCullSourceCatalog identifies a named source URL without adding bookmark or
access-lifecycle behavior.
SaliencyInfo stores only an optional subject label and confidence. It
deliberately does not import Vision or expose a Vision observation. A
PhotoAnalysisKit result can be translated into this small culling-domain value
at the host boundary.
RawParserKit’s vendor parsers return a focus string in this shape:
imageWidth imageHeight focusX focusY
FocusPointParser.normalizedPoint(from:) splits on arbitrary whitespace,
requires exactly four numeric values and positive image dimensions, and returns:
x = focusX / imageWidth
y = focusY / imageHeight
The origin remains at the visual top-left. PhotoAnalysisKit accepts that
convention and performs its own Vision coordinate conversion.
The parser does not clamp the returned point. Vendor parsers and host validation
are expected to supply meaningful sensor coordinates. Malformed input or
non-positive dimensions return nil.
5. Burst Models Preserve Evidence, Not Just A Winner
BurstAnalysisModels.swift defines the values exchanged by the two engines.
BurstGroupingOutput
├── groups: [BurstGroup]
└── boundaryEvidence: [BurstBoundaryEvidence]
BurstAnalysisResult
├── ordered candidate scores
├── recommended and second-best IDs
├── confidence and review state
├── one-click safety flag
├── reasons
└── cautions
This design avoids throwing away intermediate decisions. RawCull can show why a
boundary was created, why a file won, and why automation is or is not considered
safe.
BurstGroupingConfig.algorithmVersion is currently 4. Its defaults are visual
distance 0.25, EXIF gap 2 seconds, modification-date fallback gap 10 seconds,
same camera required, similar focal length required within 3 mm, shutter/
aperture/ISO change at most 0.5 EV, and exposure compensation at most 0.34 EV.
The version gives hosts an identity marker for persisted grouping output.
Custom decoding supplies defaults for fields that are absent from older saved
configurations.
BurstReviewState has the current workflow states .none, .needsReview,
.reviewed, and .deferred, plus compatibility states
.algorithmReviewed, .manualWinnerOverride, and .decisionApplied retained
for older caches. Unknown decoded raw values fall back to .none.
BurstWinnerOverride likewise migrates older data by generating a missing ID
and defaulting missing member filenames to an empty array.
6. BurstGroupingEngine: Decide Where A Burst Splits
The grouping engine expects files in shot order. It starts with the first file
and evaluates each adjacent pair.
flowchart TD
Pair["Previous + current file"] --> Visual{"Visual distance present\nand below threshold?"}
Visual -- No --> Split["Start new group"]
Visual -- Yes --> Time{"Time gap allowed?"}
Time -- No --> Split
Time -- Yes --> Camera{"Required camera same?"}
Camera -- No --> Split
Camera -- Yes --> Focal{"Required focal delta allowed?"}
Focal -- No --> Split
Focal -- Yes --> Exposure{"Exposure stable?"}
Exposure -- No --> Split
Exposure -- Yes --> Continue["Append to current group"]A new group begins when any enabled boundary rule fires:
- similarity evidence is missing;
- visual distance is greater than or equal to
visualDistanceThreshold; - the absolute capture gap exceeds
maxTimeGapSeconds, or
maxFallbackTimeGapSeconds when either file lacks a parsed EXIF capture
instant; - camera identity changes when
requireSameCamera is enabled; - numeric or parsed focal length changes by more than
maxFocalLengthDeltaMM
when similarity is required; - shutter speed, aperture, ISO, or exposure compensation changes by more than
its configured EV threshold.
The default EXIF capture gap is 2 seconds; the default file-date fallback gap is
10 seconds. Each boundary records captureTimeUsedFallback, allowing later
ranking to distinguish precise camera time from filesystem time.
Exposure comparison works in photographic stops:
shutter delta = |log2(current seconds / previous seconds)|
aperture delta = 2 × |log2(current f-number / previous f-number)|
ISO delta = |log2(current ISO / previous ISO)|
compensation = |current EV - previous EV|
The defaults are 0.5 EV for shutter, aperture, and ISO and 0.34 EV for exposure
compensation. The engine prefers numeric values. When a numeric shutter,
aperture, or ISO comparison is unavailable but both normalized display strings
are present and differ, it still treats the exposure as changed. The largest
available adjustment is retained as exposureAdjustmentEV.
Focal length similarly prefers focalLengthMM and falls back to the first
number in the display string. If either side has no usable value, that rule has
no delta to evaluate.
Lens changes are recorded in BurstBoundaryEvidence, but do not by themselves
split a group. They later make the group’s ranking metadata unstable. Keeping
evidence separate from the grouping decision makes this behavior visible rather
than implicit.
Missing similarity is conservative: it creates a boundary instead of assuming
two files belong together. BurstPairKey.cacheKey standardizes the ordered
adjacent-pair key used to supply those distances.
Each boundary record stores the measured values, boolean changes, final
decision, and human-readable reasons. The output assigns stable sequential group
IDs beginning at zero for that run.
7. BurstRankingEngine: Turn Evidence Into A Recommendation
Ranking consumes groups, files keyed by ID, sharpness scores, a score
normalization maximum, saliency summaries, boundary evidence, and optional
review states.
7.1 Determine Group Conditions
For each group, the engine derives:
- metadata stability: none of its internal boundaries report exposure,
camera, or lens changes;
- capture-time reliability: every member has a parsed EXIF capture date
rather than the modification-date fallback;
- tight similarity: every internal boundary has a visual distance below
0.22;
- dominant subject: the most frequent non-nil saliency label;
- burst-relative sharpness: a within-group
0...1 normalization when at
least two valid scores exist and their normalized spread is at least 0.03.
The configured grouping threshold decides membership; the fixed tighter 0.22
check contributes to ranking confidence. They answer different questions.
7.2 Score Each Candidate
If burst-relative sharpness is available, the ranking sharpness component is:
ranking sharpness = 0.65 × catalog-normalized sharpness
+ 0.35 × burst-relative sharpness
The final candidate score is:
overall = 0.62 × ranking sharpness
+ 0.12 × focus-point evidence
+ 0.10 × saliency consistency
+ 0.16 × metadata evidence
The supporting components are deliberately simple and inspectable:
| Component | Rule |
|---|
| Focus point | 0.70 when AF metadata exists, otherwise 0.45 |
| Saliency | 0.75 for the dominant label, 0.25 for a different label, 0.45 when absent |
| Metadata base | 0.70 when stable, otherwise 0.40 |
| Tight-similarity adjustment | +0.15 |
| ISO adjustment | Above ISO 1600, -0.05 per stop, capped at -0.15 |
| Wide-aperture adjustment | +0.05 at f/5.6 or wider |
| Motion-risk adjustment | +0.05 for a clearly fast shutter; up to -0.15 for a slower shutter |
Every component is clamped or normalized into a predictable range before use.
Candidates are sorted by descending overall score; equal scores preserve the
group’s original shot order.
Motion risk uses exposureTimeSeconds with focalLengthMM when both are
available. A shutter at least twice as fast as the reciprocal focal-length rule
gets the positive adjustment; a shutter slower than the reciprocal rule gets a
stop-based penalty. Without focal length, 1/500 second or faster is treated as
lower risk and 1/60 second or slower as elevated risk.
Each candidate also carries reasons and cautions such as measured sharpness,
available AF evidence, classified subject, missing sharpness, changed metadata,
fast or slow shutter behavior, and high ISO.
7.3 Assign Confidence Separately
Winning a group does not automatically mean a decision is safe to automate.
High confidence requires:
- at least one sharpness score;
- at least three files in the group;
- a best-versus-second score gap of at least 0.12;
- best absolute sharpness of at least 0.65 after normalization;
- stable metadata;
- tight visual similarity;
- parsed capture times for every member.
Medium confidence requires a gap of at least 0.05 and stable metadata. Other
results are low confidence.
isSafeForOneClickCulling is true only for high confidence.
canApplyOneClickCulling(hasSharpnessScores:) adds an explicit host-supplied
confirmation that sharpness data is available before an action is enabled.
RawCull still owns the action itself.
Reasons and cautions on BurstAnalysisResult are capped to three each so the
result remains concise enough for inspection UI and persistence.
A modification-date fallback therefore does not prevent grouping, but it does
prevent high-confidence automation and adds a capture-time caution.
8. Histogram Calculation Is Intentionally Lightweight
HistogramCalculator.normalizedLuminanceHistogram(from:) directly reads an
8-bit RGB or RGBA-style CGImage buffer. Each pixel is placed in one of 256
bins using Rec. 601 luminance:
Y = 0.299R + 0.587G + 0.114B
The bins are normalized by the largest bin count, so the peak has value 1. This
is shape normalization, not probability normalization; the bins do not
necessarily sum to 1.
The implementation requires positive dimensions, provider data, 8 bits per
component, and at least three bytes per pixel. Unsupported input returns a
zero-filled 256-bin array, preserving a stable return shape for views and
callers.
The helper is appropriate for display and lightweight comparisons. Color
conversion, RAW development, high-bit-depth analysis, and channel-layout
generalization are outside this package.
9. Concurrency Is Achieved Through Purity
RawCullCore contains no actor because it owns no shared mutable runtime state.
Models are Sendable values and engines are namespaces of nonisolated static
functions.
The package performs no asynchronous I/O. Callers do expensive decoding, Vision
work, embeddings, and sharpness analysis on the appropriate tasks, then pass
snapshots into RawCullCore.
This design has several practical benefits:
- grouping and ranking are deterministic for the same inputs;
- no application singleton can change a result midway through a call;
- tests do not require async setup or framework assets;
- background work does not cross the main actor merely because the host uses
main-actor default isolation.
10. Testing Rules And Migrations
Tests/RawCullCoreTests/ uses Swift Testing and synthetic values. Coverage
includes:
- model coding, hashing, identity, and source-catalog values;
- valid, decimal, whitespace-varied, and malformed focus strings;
- empty and stable groups;
- boundaries caused by missing distance, time, camera, focal length, and
exposure;
- EXIF capture ordering, time-zone offsets, modification-date fallback, and the
different fallback gap;
- numeric and display-string exposure comparisons in photographic stops;
- boundary evidence content;
- absolute and burst-relative sharpness ranking, motion risk, and progressive
ISO penalties;
- missing scores, stable tie order, confidence, review-state propagation, and
one-click eligibility;
- RGB and RGBA histogram bins, peak normalization, and unsupported input.
No test needs an actual catalog, RAW file, Vision model, Metal device, cache
directory, or settings store. That is the strongest evidence that the package
boundary is doing useful work.
11. How To Extend Culling Logic
Use these constraints when adding a rule:
- Pass the required fact into a
Sendable package value; do not reach into a
view model. - Preserve raw evidence separately from the final boolean or winner.
- Keep scoring weights and confidence thresholds deterministic and testable.
- Decide whether a rule affects group membership, candidate ranking,
confidence, or only a caution.
- If evidence changes a grouping boundary or its meaning, bump
BurstGroupingConfig.algorithmVersion and add migration/engine tests. - If the app’s persisted result shape, validation, or pipeline identity changes,
bump
BurstAnalysisCache.schemaVersion (currently 9) and update cache tests. - Maintain decoding fallbacks for existing review and override data.
- Leave framework observations and heavy image processing in the producing
package.
- Leave user actions and presentation state in RawCull.
Source Map
| Topic | RawCullCore source |
|---|
| Product and concurrency settings | Package.swift |
| EXIF transport value | Sources/RawCullCore/ExifMetadata.swift |
| File and source identity | Sources/RawCullCore/RawCullFileItem.swift, RawCullSourceCatalog.swift |
| Saliency summary | Sources/RawCullCore/SaliencyInfo.swift |
| Focus normalization | Sources/RawCullCore/FocusPointParser.swift |
| Burst contracts and migrations | Sources/RawCullCore/BurstAnalysisModels.swift |
| Boundary decisions | Sources/RawCullCore/BurstGroupingEngine.swift |
| Ranking and confidence | Sources/RawCullCore/BurstRankingEngine.swift |
| Histogram | Sources/RawCullCore/HistogramCalculator.swift |
| Behavior tests | Tests/RawCullCoreTests/ |
| RawCull metadata adapter and burst orchestration | RawCull/Model/RawImageLoading.swift, RawCull/Model/ViewModels/RawCullViewModel+BurstGrouping.swift |
Return to the package overview in RawCull Packages, or
continue with the file-decoding layer in
How RawParserKit Is Constructed.
11.4 - How RawParserKit Is Constructed
A detailed guide to RawParserKit’s vendor dispatch, TIFF and MakerNote parsing, embedded previews, structured capture and exposure metadata, orientation, decode limiting, cancellation, compatibility APIs, and tests.
How RawParserKit Is Constructed
Revision scope: This walkthrough was written against RawParserKit 1.3.0
at d2175ed880d39021bdb5f5a2a842b460af0b316c. The current RawCull
checkout resolves 1.3.1 at f0e5b02a10294798afd86781da5d6510e146a7ba.
The detailed examples below document the earlier reviewed source; check the
current package revision when changing parser behavior.
RawParserKit is RawCull’s camera-file boundary. It knows how Sony ARW, Nikon
NEF, and Adobe DNG files are structured, how to locate their embedded JPEGs and
AF metadata, and how to turn those sources into orientation-normalized images
and display-ready metadata.
The package stops at decoding. It does not score sharpness, generate embeddings,
group bursts, cache application results, or decide which photo should be kept.
1. Begin With The Package Boundary
RawParserKit owns:
- vendor-neutral RAW format dispatch;
- Sony ARW, Nikon NEF, and Adobe DNG format knowledge;
- TIFF IFD and vendor MakerNote traversal;
- focus-location and embedded-JPEG offset parsing;
- RAW and rendered-image thumbnails and previews;
- source-orientation normalization;
- display-ready and numeric EXIF/RAW metadata extraction, including a structured
capture instant and time-zone offset;
- Sony full sensor development to JPEG;
- decode task deduplication, concurrency limiting, and cooperative cancellation;
- diagnostics and staged compatibility APIs.
The host application owns:
- security-scoped URL access and sandbox bookmarks;
- directory discovery, catalogs, and source selection;
- memory and disk cache locations and eviction policy;
- image-analysis inputs and result identity;
- observable loading state, placeholders, retries, and presentation;
- ratings, burst grouping, ranking, and saved-file behavior.
RawParserKit supplies inputs to higher layers without importing them:
flowchart LR
File["ARW, NEF, DNG, JPEG, PNG, or TIFF"] --> Parser["RawParserKit"]
Parser --> Image["CGImage or NSImage"]
Parser --> Metadata["RawImageMetadata + RawFocusPoint"]
Image --> Host["RawCull adapters"]
Metadata --> Host
Host --> Analysis["PhotoAnalysisKit / PhotoAIKit"]
Host --> Core["RawCullCore"]2. Package Shape And Framework Boundary
Package.swift declares Swift tools 6.2, Swift 6 language mode, macOS 26, one
library product, and one test target. There are no third-party package
dependencies.
The implementation uses Apple frameworks appropriate to file decoding:
Foundation, AppKit, Core Graphics, ImageIO, Core Image, OSLog, and
synchronization primitives from os.
Like RawCullCore, the target enables main-actor default isolation plus
InferIsolatedConformances and NonisolatedNonsendingByDefault. Pure parsing
and static format APIs opt out with nonisolated; the stateful loader and
limiter use actors.
The source is arranged in layers:
RawImageLoader high-level deduplicated facade
├── RawFormatRegistry + RawFormat vendor-neutral dispatch
│ ├── SonyRawFormat
│ ├── NikonRawFormat
│ └── DNGRawFormat
├── thumbnail and preview extractors ImageIO + binary fallback
├── MakerNote parsers TIFF byte traversal
├── OrientationNormalizedImageLoader rendered/embedded image helpers
└── cancellation + decode limiter concurrency control
Callers can use the facade for normal browser behavior or a lower layer when
they need explicit control.
Sources/RawParserKit/RawFormat.swift describes the static capabilities every
camera format supplies:
- supported filename extensions and a display name;
- thumbnail extraction;
- embedded-preview extraction;
- AF focus-location parsing;
- a readable label for compression codes;
- camera-specific megapixel thresholds for S, M, and L size classes.
Conformers are stateless enums. RawFormatRegistry.all currently registers:
| Conformer | Extension | Vendor responsibilities |
|---|
SonyRawFormat | .arw | Sony extractors, MakerNote parser, compression labels, body thresholds, and full sensor JPEG creation |
NikonRawFormat | .nef | Nikon extractors, MakerNote parser, compression labels, and body thresholds |
DNGRawFormat | .dng | DNG TIFF/SubIFD parser, extractors, compression labels, and generic/camera-family thresholds |
format(for:) lowercases a URL’s extension and returns a format metatype.
Callers invoke static protocol requirements on that value without switching on
brands.
The default rawSizeClass implementation converts dimensions to megapixels,
obtains body-specific L and M thresholds, and returns L, M, or S. Unknown
bodies use a generic fallback supplied by the conformer.
This is a simple plugin architecture inside the package. Adding a camera brand
means adding a conformer and registering it, not adding vendor switches
throughout the loader.
4. RawImageLoader Is The High-Level Facade
RawImageLoader.shared is an actor. It offers four current operations:
| API | Result |
|---|
thumbnail(for:maxPixelSize:) | An NSImage suitable for grids and browsers |
thumbnailCGImage(for:maxPixelSize:) | The same thumbnail as CGImage |
previewImage(for:) | A larger sidecar or embedded preview |
metadata(for:) | A neutral RawImageMetadata snapshot |
4.1 Deduplicate Equivalent Work
The actor stores in-flight tasks:
- thumbnails keyed by URL and requested pixel size;
- previews keyed by URL;
- metadata keyed by URL.
If another caller requests the same work while it is running, both await the
existing task. The entry is removed after completion. This prevents rapid view
updates from decoding the same large image repeatedly.
4.2 Bound Expensive Decodes
The facade uses two DecodeConcurrencyLimiter actors:
- up to six concurrent thumbnail decodes;
- up to two concurrent full-size preview decodes.
These limits control memory as much as CPU. A few full-resolution bitmaps can
consume substantially more memory than the compressed RAW files that produced
them.
4.3 Follow The Thumbnail Strategy
For a rendered JPEG, PNG, or TIFF, the loader asks ImageIO for an
orientation-aware thumbnail. For RAW input, it tries an embedded ImageIO
thumbnail first and then dispatches to the registered vendor extractor. The
vendor result is orientation-normalized before it becomes an NSImage.
4.4 Follow The Preview Strategy
The preview path checks for a same-basename .jpg sidecar first. If absent, it
tries an orientation-aware embedded preview and then the registered format’s
embedded-preview extractor. Vendor output is normalized using the source
orientation.
The sidecar-first decision is host-facing convenience, not RAW parsing. A caller
that requires only bytes physically embedded in the RAW can call the format or
extractor API directly.
RawImageMetadata contains optional display and numeric values for camera,
lens, shutter speed, exposure time in seconds, aperture, focal length in
millimeters, ISO, exposure compensation in EV, capture time, dimensions, focus
point, RAW compression label, size class, and pixel dimensions.
RawImageLoader.metadata(for:) combines several sources:
- ImageIO properties from the source or a sidecar fallback;
- TIFF make, model, compression, and display-date fields;
- EXIF exposure time, aperture, focal length, ISO, exposure bias, lens, and
dimensions;
- EXIF
DateTimeOriginal, SubsecTimeOriginal, and OffsetTimeOriginal; - the registered vendor MakerNote parser for an AF point;
- EXIF subject-area coordinates when vendor focus data is unavailable;
- vendor-specific compression labels and size-class thresholds.
captureDate parses the original capture timestamp as a real Date, preserving
variable-length fractional seconds. OffsetTimeOriginal accepts Z, +HH:MM,
-HH:MM, and compact +HHMM/-HHMM forms; the parsed offset is also retained
as captureTimeZoneOffsetSeconds. When the EXIF offset is missing, parsing uses
the current time zone. A missing or invalid DateTimeOriginal leaves
captureDate nil rather than substituting the TIFF display date.
capturedAt remains a display string and can fall back to TIFF DateTime.
Keeping it separate from captureDate prevents formatted UI text from becoming
burst-ordering evidence.
rows returns non-empty display label/value pairs, while isEmpty lets the
facade omit an empty metadata object. Numeric exposure time, aperture, focal
length, ISO, exposure compensation, width, and height remain available so hosts
do not have to parse formatted strings for analysis or domain rules.
RawFocusPoint stores normalized X and Y coordinates. Its failable focus-string
initializer accepts the common "width height x y" shape, checks positive
dimensions, and rejects coordinates outside 0...1.
RawCull’s RawParserKitImageLoader maps this parser-owned snapshot into its own
ExifMetadata and RawCullFileItem values. ScanFiles carries the parsed
capture instant and offset into the catalog; RawCullCore can then prefer camera
time and explicitly detect a modification-date fallback. The similar types exist
on purpose: each package owns the vocabulary at its boundary and neither must
depend on the other.
Thumbnails should not require full RAW development when a camera already stored
a usable JPEG.
6.1 Sony
SonyThumbnailExtractor first asks SonyMakerNoteParser for embedded JPEG
locations and decodes a selected JPEG directly. This bypasses macOS RAW decoder
failures seen with newer ARW layouts. If the binary path cannot locate a JPEG,
it falls back to ImageIO’s embedded-thumbnail behavior.
6.2 Nikon
NikonThumbnailExtractor asks ImageIO for a transformed embedded thumbnail.
Both vendor extractors then redraw into an 8-bit premultiplied sRGB bitmap using
interpolation quality derived from qualityCost.
6.3 DNG
DNGThumbnailExtractor follows the same cancellation-aware contract. Its
binary fallback uses DNGMakerNoteParser and TIFF IFD/SubIFD classification to
avoid treating JPEG-compressed raw image data as a display preview.
Both APIs run their synchronous ImageIO work through CancellableImageIOWork
and throw ThumbnailError for an invalid source, failed generation, or failed
bitmap context.
ThumbnailSharpener is an optional lower-level Core Image helper for producing
a sharpened preview at a requested maximum dimension. The high-level boundary
does not force sharpening on every caller.
The three formats use the same building blocks in a different order.
SonyEmbeddedJPEGExtractor prefers the binary TIFF locator, avoiding RAW
decoder initialization on affected ARW files, then falls back to ImageIO.
NikonEmbeddedJPEGExtractor inspects ImageIO sub-images first and uses the
binary locator when the preview is not exposed there.
flowchart TD
Source["RAW URL"] --> ImageIO["Inspect ImageIO image indexes"]
Source --> Parser["Vendor TIFF parser"]
ImageIO -->|usable JPEG| Decode["Decode or downsample"]
Parser --> Offset["Absolute JPEG offset + length"]
Offset --> Bytes["Read JPEG bytes"]
Bytes --> Decode
Decode --> Result["CGImage"]fullSize: true permits a longest edge up to 8640 pixels. The normal preview
path limits large images to 4320 pixels.
The Nikon fallback prefers the full-resolution SubIFD preview for full-size
requests and IFD1 for smaller requests. Sony chooses the largest available JPEG
first, with preview and thumbnail fallbacks.
DNG prefers standards-classified preview IFDs using NewSubFileType and
Compression; files that omit NewSubFileType retain the positional fallback
needed by older or nonconforming writers.
Extractor-level limiters default to two concurrent operations, and a caller can
inject a shared limiter. This allows a host facade to enforce one budget across
several decode paths instead of accidentally stacking independent limits.
SonyRawFormat.createFullSizeJPEG(from:quality:) does not return a
camera-embedded preview. SonyRAWJPEGCreator develops the ARW sensor data
through macOS CIRAWFilter, renders it into sRGB, and encodes JPEG data.
Quality must be in 0...1. The operation can fail with:
invalidQuality;unsupportedOrInvalidRAW when the installed macOS RAW decoder cannot develop
that file;encodingFailed.
This distinction matters in UI and caching. An embedded preview and a developed
sensor image have different cost, pixels, appearance, and invalidation semantics
even if both are JPEG-encoded at the end.
9. Sony MakerNote And TIFF Parsing
Sony ARW is TIFF-based. Focus parsing follows this structure:
TIFF IFD0
└── ExifIFD tag 0x8769
└── MakerNote tag 0x927C
└── Sony MakerNote IFD
└── FocusLocation tag 0x2027
(fallback: 0x204A)
The focus tag contains four unsigned 16-bit values: image width, image height,
X, and Y. Sony MakerNote IFD offsets are interpreted as absolute file offsets.
The production parser uses a fast path and a fallback:
- focus location reads the first 4 MB, then retries with the full file when
necessary;
- embedded-JPEG discovery reads the first 512 KB, then retries the full file
when no locations were found.
The embedded locator walks TIFF IFDs and returns optional absolute locations for
a small thumbnail, preview, and full JPEG. readEmbeddedJPEGData seeks directly
to a validated location and reads its byte range.
The fast path keeps normal scans inexpensive; the full-file fallback supports
bodies that place relevant TIFF structures near the end of the file.
10. Nikon MakerNote And TIFF Parsing
Nikon NEF is also TIFF-based, but modern Nikon Type-3 MakerNotes contain their
own TIFF header:
TIFF IFD0
└── ExifIFD tag 0x8769
└── MakerNote tag 0x927C
├── "Nikon\0" signature + version
└── inner TIFF header
└── Nikon IFD
└── AFInfo2 tag 0x00B7
Offsets in the inner TIFF are relative to the MakerNote TIFF-header base, not
the start of the NEF file. Keeping this offset rule inside the Nikon parser
prevents a generic loader from acquiring vendor-specific exceptions.
For supported modern AFInfo2 layouts, the parser reads AF image dimensions, area
position, and area size, and returns the same "width height x y" shape as
Sony. The public shape lets all downstream code use one focus-point adapter.
Nikon embedded preview discovery examines Compression=6 SubIFDs referenced from
IFD0 and the IFD1 JPEG interchange fields. It returns optional locations for the
largest preview and IFD1 JPEG.
As on Sony, focus parsing starts with 4 MB and falls back to the full file.
Embedded-location parsing uses a 1 MB fast path followed by a full-file retry
when required.
10.1 DNG TIFF And SubIFD Parsing
DNG uses the same neutral focus-location string but has no single camera-vendor
MakerNote layout. DNGMakerNoteParser walks TIFF IFD0 and SubIFDs, uses standard
EXIF focus evidence when present, and exposes DNGEmbeddedJPEGLocations with
thumbnail, preview, and full-JPEG candidates. Standards-classified previews are
selected from TIFF NewSubFileType and Compression values; positional rules
are used only when the classification tag is absent.
DNGRawFormat reports container-appropriate compression names for uncompressed,
JPEG, Deflate, PackBits, Lossy DNG, and JPEG XL values. Its size-class policy is
megapixel-based with small camera-family overrides because DNG is a cross-vendor
container rather than a single body line.
11. Diagnostics Report The Failed Stage
The ordinary parser APIs return optionals because missing or unsupported
MakerNote data is expected during normal browsing.
For troubleshooting, each vendor also exposes diagnostic forms for focus and
embedded-JPEG lookup. RawParserDiagnostics<Value> contains:
- the optional parsed value;
- an ordered trace of stages and offsets checked;
- an optional final failure explanation.
This keeps logging policy outside the binary parser while allowing RawCull’s
diagnostics UI to show whether failure occurred at file access, TIFF validation,
IFD lookup, MakerNote traversal, tag interpretation, or fallback.
12. Orientation Helpers Normalize Visual Coordinates
OrientationNormalizedImageLoader provides lower-level operations for rendered
URLs, encoded JPEG data, source thumbnails, embedded thumbnails, and embedded
previews.
ImageIO’s transform option is used when available. The loader also implements
the eight EXIF orientation transforms, including mirrored and transposed cases,
and can read orientation from the RAW source when decoding separately extracted
JPEG data.
SupportedFileType enumerates .arw, .nef, .jpeg, .jpg, .png, .tif,
and .tiff. Its rendered-image set distinguishes sources that can be loaded
directly from formats that need RAW dispatch.
Orientation is part of the parser boundary because AF coordinates and subject
analysis must refer to the same visual image the user sees.
13. Cancellation And Decode Limiting Solve Different Problems
CancellableImageIOWork bridges a synchronous ImageIO closure to async code on
a global dispatch queue. It creates an ImageIOCancellationToken and uses a
locked state machine to ensure the checked continuation is resumed exactly once,
even when cancellation races completion.
Cancellation is cooperative. The token is checked before and after synchronous
framework calls and between multi-stage loops. A framework function already
executing may not stop internally, but its result is discarded when cancellation
is observed.
DecodeConcurrencyLimiter is an actor that solves admission control. It:
- grants work immediately while slots are available;
- queues additional continuations;
- removes and resumes a cancelled waiter;
- transfers a released slot directly to the next waiter;
- releases the slot with
defer when work completes.
Cancellation prevents obsolete work; limiting prevents too much valid work from
running simultaneously. The loader needs both.
14. RawCull Adapter And Image Ownership
RawCull depends on its own RawImageLoading: Sendable protocol. The
RawParserKitImageLoader value forwards to RawImageLoader.shared and performs
the package-to-app translation:
| RawCull request | Package facade call | Boundary result |
|---|
fileMetadata(for:) | metadata(for:) | RawImageMetadata becomes app ExifMetadata, capture date/offset, legacy focus string, and normalized CGPoint |
thumbnailCGImage(for:maxPixelSize:) | thumbnailCGImage(for:maxPixelSize:) | CGImage for cache and analysis paths |
thumbnailImage(for:maxPixelSize:) | thumbnail(for:maxPixelSize:) | NSImage consumed by app/UI-isolated code |
previewCGImage(for:) | previewImage(for:) | CGImage for the preview and analysis pipeline |
App code should enter through this facade or the vendor-neutral registry. Sony
and Nikon conformers remain implementation details except in diagnostics and
package tests.
NSImage and CGImage are framework reference objects rather than ordinary
Sendable values. RawCull therefore keeps their lifetime inside the actor or UI
operation that needs them. RequestThumbnail converts a CGImage to JPEG
Data inside its actor before starting the detached disk save; the value
crossing that task boundary is Data, not the image object. Prefer the
CGImage facade methods for background analysis and convert to presentation
objects at the presentation boundary.
15. Compatibility APIs Support Staged Migration
The package retains deprecated names such as:
BrowserExifInfo and BrowserFocusPoint;thumbnail200px, extractembeddedJPG, and exifInfo;JPGSonyARWExtractor and JPGNikonNEFExtractor;extractFullJPEG on RawFormat.
Each shim forwards to a current neutral name. This lets RawCull migrate call
sites without forcing an all-at-once source break, while deprecation warnings
make the intended direction visible.
New code should use RawImageMetadata, RawFocusPoint, thumbnail,
previewImage, metadata, the embedded-JPEG extractors, and
extractEmbeddedPreview.
16. Test Binary Rules Without Shipping Camera Files
Tests/RawParserKitTests/ uses Swift Testing and mostly synthetic TIFF-like
byte buffers. The suite covers:
- registry extension matching and dispatch;
- Sony focus tags, offset rules, invalid byte-order markers, fallback tags,
embedded JPEG locations, and diagnostics;
- Nikon Type-3 MakerNote and AFInfo2 layouts, SubIFDs, IFD1 JPEGs, offset rules,
and diagnostics;
- DNG TIFF/SubIFD classification, focus and embedded-JPEG locations,
compression labels, size classes, and malformed-data behavior;
- direct reading of embedded JPEG bytes;
- cancellation before decode and cancellation behavior in vendor extractors;
- decode-limiter capacity;
- Sony full-size JPEG quality validation and generated JPEG properties;
- numeric exposure-time, focal-length, aperture, ISO, and exposure-compensation
metadata;
- capture-date offsets, subsecond precision, invalid-date behavior, and offset
parsing;
- current public naming and format-helper behavior.
Synthetic binary fixtures make edge cases reproducible and avoid committing
large proprietary ARW, NEF, and DNG samples. Framework integration is tested with
small generated images where needed.
17. How To Add Another Camera Vendor
Use the existing extension points:
- Add a stateless
RawFormat conformer with its extension and display name. - Implement vendor-specific thumbnail, preview, focus, compression, and
size-class behavior.
- Put TIFF or MakerNote byte rules in a dedicated parser, including explicit
offset bases and bounds checks.
- Return the common normalized focus-location string at the public boundary.
- Add ImageIO behavior first and a binary embedded-JPEG fallback when the
framework does not expose the preview reliably.
- Make blocking decode stages cooperative with
CancellableImageIOWork. - Register the conformer in
RawFormatRegistry.all. - Add synthetic binary tests for endian, offsets, missing tags, corrupt
lengths, and diagnostics.
- Leave analysis, cache placement, and UI policy in their owning layers.
Source Map
| Topic | RawParserKit source |
|---|
| Product and concurrency settings | Package.swift |
| Format contract and dispatch | Sources/RawParserKit/RawFormat.swift, RawFormatRegistry.swift |
| Format conformers | Sources/RawParserKit/SonyRawFormat.swift, NikonRawFormat.swift, DNGRawFormat.swift |
| High-level facade | Sources/RawParserKit/RawImageLoader.swift |
| Metadata and focus values | Sources/RawParserKit/BrowserExifInfo.swift, BrowserFocusPoint.swift |
| Sony TIFF and MakerNote parsing | Sources/RawParserKit/SonyMakerNoteParser.swift |
| Nikon TIFF and MakerNote parsing | Sources/RawParserKit/NikonMakerNoteParser.swift |
| DNG TIFF and preview parsing | Sources/RawParserKit/DNGMakerNoteParser.swift |
| Embedded preview extraction | JPGSonyARWExtractor.swift, JPGNikonNEFExtractor.swift, DNEmbeddedJPEGExtractor.swift |
| Thumbnail extraction | SonyThumbnailExtractor.swift, NikonThumbnailExtractor.swift, DNGThumbnailExtractor.swift |
| Full Sony sensor development | Sources/RawParserKit/SonyRAWJPEGCreator.swift |
| Orientation and rendered files | Sources/RawParserKit/OrientationNormalizedImageLoader.swift |
| Cancellation bridge | Sources/RawParserKit/CancellableImageIOWork.swift |
| Decode admission control | Sources/RawParserKit/DecodeConcurrencyLimiter.swift |
| Parser diagnostics | Sources/RawParserKit/RawParserDiagnostics.swift |
| Behavior and binary-fixture tests | Tests/RawParserKitTests/ |
| RawCull metadata adapter and catalog scan | RawCull/Model/RawImageLoading.swift, RawCull/Actors/ScanFiles.swift |
Continue from decoded images into
How PhotoAnalysisKit Is Constructed, or return to the
package overview.
12 - File Read and Write Reference
Files, folders, and persistent data touched by RawCull
File Read and Write Reference
This page lists the main places RawCull reads and writes files. Use it before
changing sandbox access, cache locations, persistence, or export behavior.
File Map
| File/folder | Access | Owner |
|---|
| User-selected catalog folder | Read | RawCullViewModel, ScanFiles, DiscoverFiles, parser package |
RAW files (.arw, .nef, .dng) | Read | scan, thumbnails, focus parsing, zoom, export, diagnostics |
focuspoints.json beside catalog | Read optional | ScanFiles fallback |
App Support savedfiles.json and backups | Read/write/move | CullingModel, ReadSavedFilesJSON, WriteSavedFilesJSON |
App Support settings.json | Read/write | SettingsViewModel, SettingsFileWriter |
| App Support analysis artifacts and burst snapshots | Read/write/delete | PerFileAnalysisArtifactStore, BurstAnalysisCache |
| App Support AI models and licence acceptance | Read/write | AI model download/resource and licence services |
| Thumbnail cache directory | Read/write/delete | DiskCacheManager |
| Full-size JPEG preview cache | Read/write/prune | FullSizeJPGDiskCache, ZoomPreviewHandler |
| Subject-mask cache | Read/write/delete | PhotoAIKit subject-mask stores configured by RawCullAIIntegration |
Exported .jpg files in a chosen destination | Write | ExtractAndSaveJPGs, SaveJPGImage |
| Temporary rsync include lists / process streams | Write/delete/read | ExecuteCopyFiles, ArgumentsSynchronize, PrepareOutputFromRsync |
| Destination security-scoped bookmark | Read/write UserDefaults | OpencatalogView, copy workflow |
| AI selections and managed-model metadata | Read/write UserDefaults and app metadata | RawCullAISettingsModel, model download service |
Catalog Reads
The active catalog comes from the sidebar folder selection.
RawCullViewModel.startCatalogLoad(for:) starts security-scoped access and then
runs the scan.
Catalog reads include:
- directory enumeration,
- URL resource values,
- EXIF metadata via ImageIO,
- MakerNote focus points via
RawParserKit, - embedded thumbnails/JPEGs,
- optional
focuspoints.json.
DiscoverFiles uses RawFormatRegistry.allExtensions so it follows the parser
registry.
App Support Files
Application Support is used for durable app-owned data:
~/Library/Application Support/RawCull/
Important files:
| File | Purpose |
|---|
savedfiles.json | Ratings, sharpness/saliency persistence, and manual burst winner overrides |
savedfiles.backup.json | Atomic backup of the previous valid saved-file store before replacement |
savedfiles-corrupt-<timestamp>.json | User-approved archive of a store that failed decoding |
settings.json | Thumbnail, cache, scoring, and focus-mask settings |
AnalysisArtifacts/ | Per-file, descriptor-valid Vision/CLIP similarity artifacts |
BurstAnalysis/ | Derived catalog snapshots containing grouping, ranking, artifacts, and review states |
Models/ | Installed AI model bundles grouped by model identity |
ModelLicenceAcceptances.json | Recorded model-licence acceptance state |
CopyLists/ | Operation-unique NUL-separated rsync include lists, removed during cleanup |
savedfiles.json is written atomically after the old data is copied atomically
to savedfiles.backup.json. A decode failure is surfaced to the UI; rating
mutations are blocked until the user retries or explicitly archives the damaged
store. settings.json, per-file artifacts, and burst snapshots also use atomic
replacement. Burst-analysis validity is checked against file metadata,
descriptors, artifact digest, and algorithm/signature versions before reuse.
Cache Files
Generated caches live under the user cache directory for the RawCull app
identifier. They are performance data, not source-of-truth data.
| Cache | Purpose |
|---|
| Schema-specific thumbnail disk cache | Stores generated JPEG representations keyed by source fingerprint, purpose, requested size, and orientation policy |
| Full-size JPEG disk cache | Stores larger embedded JPEG previews for zoom |
| Subject-mask cache | Stores reusable segmentation masks outside the durable app-data namespace |
Deleting these caches should only make RawCull slower until they are rebuilt. It
should not lose ratings or manual decisions.
Exported JPEGs
ExtractAndSaveJPGs exports the current selection into a user-selected
destination catalog. It supports two modes:
| Mode | Input path | Output name |
|---|
| Embedded JPG | FullSizePreviewLoader.loadEmbeddedPreview | Original basename plus .jpg |
| Demosaiced RAW | SonyRawFormat.createFullSizeJPEG | Original basename plus _demosaic.jpg |
The actor bounds parallel extraction, tracks progress and per-file failures, and
passes JPEG Data to SaveJPGImage; non-Sendable image objects do not cross
the save boundary. SaveJPGImage creates files without overwriting. If a name
already exists, it retries with (1), (2), and so on; the filesystem
enforces exclusivity for case-insensitive and simultaneous-export collisions.
RawCullViewModel.startSelectedJPGExtraction starts destination security-scoped
access before constructing the actor and stops it on the main actor after the
awaited result returns.
rsync Copy Workflow
The copy workflow is separate from thumbnail/scoring export. It uses rsync to
copy selected RAW files based on rating/tag choices.
Main files:
| File | Role |
|---|
CopyFilesView.swift | UI and execution lifecycle |
OpencatalogView.swift | Destination picker and bookmark creation |
ExecuteCopyFiles.swift | Process owner and progress/result state |
ArgumentsSynchronize.swift | Builds rsync arguments |
PrepareOutputFromRsync.swift | Parses process output |
RemoteDataNumbers.swift | Summarizes copied file counts and sizes |
ExecuteCopyFiles.startcopyfiles first derives the selected filenames from the
current RawCullViewModel. It then creates an operation-unique file under
Application Support/RawCull/CopyLists/. Each UTF-8 filename is terminated by
NUL, and rsync receives --from0 plus --files-from=<path>. This preserves
spaces and newlines without converting the list into command-line arguments.
The source is the currently selected catalog URL and the destination is restored
only from destBookmark; there is no arbitrary path fallback. Both successful
scope acquisitions remain owned by the ExecuteCopyFiles instance while
/usr/bin/rsync runs. A stale destination bookmark is refreshed while its
resolved grant is active. Process handlers stream progress and a typed
CopyOutcome (success, failed, or cancelled) back to main-actor state.
Startup returns a typed CopyStartupFailure for unavailable arguments, missing
model state, an empty selection, Application Support/include-list failures,
security-scope failures, and process-launch failures. All failure paths call the
same idempotent cleanup used by completion, cancellation, close(), and
deinitialization. Cleanup finishes the progress stream, stops both acquired
scopes exactly once, removes only this operation’s include-list file, and
releases process handlers.
Security-Scoped Bookmarks
The copy workflow stores destination bookmark Data in UserDefaults after the
user picks a folder. Picker access is balanced immediately after bookmark
creation. At execution time, ExecuteCopyFiles starts a fresh scope for the
active catalog URL and resolves destBookmark with .withSecurityScope for the
operation-lifetime destination scope. Failure asks the user to reopen the
catalog or reselect the destination rather than attempting a plain-path fallback.
The catalog browsing flow is different: RawCullViewModel owns one active
security-scoped catalog URL and stops it during catalog transition or successful
application termination. Do not transfer that ownership implicitly to a child
actor.
See Security-Scoped URLs for lifecycle details.
Settings
SettingsViewModel stores its SavedSettings value as pretty-printed, sorted
JSON in Application Support/RawCull/settings.json, not in UserDefaults.
Settings affect:
- thumbnail sizes,
- cache size maximums,
- focus/scoring options,
- memory/cache defaults.
The main-actor model loads once through ensureLoaded(). Encoding happens on
the main actor from a consistent observable snapshot, and SettingsFileWriter
performs directory creation and atomic writing through an actor. Background
actors use SettingsViewModel.shared.asyncgetsettings() to obtain a Sendable
SavedSettings value rather than reading observable properties across isolation
boundaries.
AI model selection and copy bookmarks are separate preferences and may still use
UserDefaults; do not treat those as part of settings.json without an
explicit migration.
Diagnostics Reads
RawCull currently has no separate RAW-diagnostics report file or persistent
similarity-diagnostics log in the app target. Developer diagnostics use OSLog,
package tests, and focused app integration tests. If a file-backed log is added
later, keep it bounded, app-owned, and free of full user paths in ordinary
presentation.
What To Check When Changing This Area
- Writes outside the app container need active security-scoped access.
- App-owned durable data belongs in Application Support, not Caches.
- Rebuildable performance data belongs in Caches, not Application Support.
- Keep
savedfiles.json backup/corruption recovery semantics when changing
culling persistence. - Keep
settings.json separate from bookmarks and AI-selection preferences
unless a migration is designed. - If a cache stores derived algorithm output, include enough version/signature
metadata to reject stale data.
- Keep process-output parsing separate from process lifecycle management.
- Keep rsync include lists operation-unique and remove them on success, failure,
cancellation, and deinitialization.
- Balance every successful security-scope start exactly once at the layer that
owns its lifetime.
13 - Synchronous Code
Synchronous Code
Most RawCull code uses async/await, actors, and task groups. Some framework and system APIs are still synchronous: ImageIO decode, Core Image RAW rendering, JPEG encoding, filesystem calls, binary MakerNote parsing, and process launch. async on a caller does not make one of those calls nonblocking. This page records where the blocking work lives and which execution strategy each path uses.
Source Map
| Area | Files |
|---|
| Cancellation-aware blocking bridge | RawParserKit/Sources/RawParserKit/CancellableImageIOWork.swift |
| RAW thumbnail and embedded JPEG extraction | SonyThumbnailExtractor.swift, NikonThumbnailExtractor.swift, DNGThumbnailExtractor.swift, JPGSonyARWExtractor.swift, JPGNikonNEFExtractor.swift, DNEmbeddedJPEGExtractor.swift |
| RAW development and orientation | SonyRAWJPEGCreator.swift, ThumbnailSharpener.swift, OrientationNormalizedImageLoader.swift |
| Binary RAW parsing | Sony, Nikon, and DNG MakerNote/format files in RawParserKit/Sources/RawParserKit/ |
| Thumbnail and preview callers | RequestThumbnail.swift, ScanAndCreateThumbnails.swift, ScanAndExtractJPGs.swift, FullSizePreviewLoader.swift, ZoomPreviewHandler.swift, ComparisonImageLoader.swift |
| Export and JPEG encoding | ExtractAndSaveJPGs.swift, SaveJPGImage.swift, DiskCacheManager.swift, FullSizeJPGDiskCache.swift |
| Filesystem and persistence | ScanFiles.swift, DiscoverFiles.swift, PerFileAnalysisArtifactStore.swift, SettingsViewModel.swift, ReadSavedFilesJSON.swift, WriteSavedFilesJSON.swift |
| Image analysis | RawCullPhotoAnalysisAdapter.swift, DeepAIReviewFeature.swift, DeepAIReviewMaskOutlineRenderer.swift |
| External process | ExecuteCopyFiles.swift, RsyncProcessStreaming.RsyncProcess |
Why Blocking Work Matters
Swift’s cooperative thread pool expects async tasks to suspend instead of occupying threads for long periods. ImageIO and CoreImage calls often do not suspend; they block until decode/render work is done.
If RawCull runs those calls directly inside many task-group children, the calls can occupy the cooperative pool together. The UI may remain on the main actor, but unrelated async work can stop making progress and cancellation can appear late. Actor isolation prevents data races; it does not prevent a synchronous call from monopolizing the thread executing that actor.
The Bridge
CancellableImageIOWork.run(qos:_:) wraps blocking work like this:
flowchart LR
A["Swift async caller"] --> B["withTaskCancellationHandler"]
B --> C["withCheckedThrowingContinuation"]
C --> D["DispatchQueue.global(qos).async"]
D --> E["Synchronous ImageIO/CoreImage operation"]
E --> F["Resume continuation once"]The operation receives an ImageIOCancellationToken. The token checks its own lock-backed cancellation flag and Task.isCancelled. Cancellation resumes the awaiting continuation promptly, but it cannot forcibly interrupt an ImageIO or Core Image call already executing. The worker must reach a checkpoint before it observes cancellation.
WorkState protects the continuation with a lock so cancellation and completion races resume exactly once.
The parser package uses stateless enum extractors. A typical extractor exposes an async public API and a private synchronous implementation:
public async extractThumbnail(...)
-> CancellableImageIOWork.run(...)
-> private extractSync(...)
That shape appears in the Sony, Nikon, and DNG thumbnail/embedded-preview
extractors. The compatibility enums JPGSonyARWExtractor and
JPGNikonNEFExtractor remain deprecated public shims.
SonyRAWJPEGCreator.createFullSizeJPEG uses the same bridge at utility QoS for CIRAWFilter, render probing, and JPEG representation. DecodeConcurrencyLimiter separately bounds how many expensive decodes are admitted; limiting concurrency and moving blocking work off the cooperative executor solve different problems and both protections should remain.
Blocking-Work Audit
| Synchronous operation | Current execution boundary | Why |
|---|
| ARW/NEF/DNG thumbnail and embedded-preview ImageIO decode | CancellableImageIOWork on a global GCD queue, with cancellation checkpoints and decode limiting where supplied | Decode duration is input- and OS-decoder-dependent and may be repeated across a catalog. |
Sony developed RAW via CIRAWFilter and JPEG representation | CancellableImageIOWork at utility QoS | Full RAW development and encoding are long, nonsuspending framework calls. |
| Sharpened RAW preview | Detached task in ZoomPreviewHandler; concurrent task in ComparisonImageLoader | ThumbnailSharpener performs synchronous CIRAWFilter and CIContext rendering. The detached zoom path is appropriate for potentially long rendering; a concurrent task alone still uses Swift’s cooperative executor and must remain bounded. |
| Orientation-normalized ImageIO load | Detached task in disk/full-size preview cache callers; otherwise kept inside an already isolated worker | File decode can block. The synchronous loader is a leaf API, so the caller owns the execution boundary. |
| Thumbnail/full-size cache reads, writes, size scans, and pruning | Task.detached with user-initiated, background, or utility priority according to latency | Data.write, directory enumeration, resource-value reads, ImageIO cache decode, and deletion are filesystem-bound and nonsuspending. |
| JPG export writes | Encode CGImage to Sendable Data in the owning actor, then atomic Data.write in a detached background task | Avoids sending a non-Sendable image across isolation and keeps file writes off the actor/cooperative executor. |
| Catalog enumeration | Synchronous contentsOfDirectory at the start of the ScanFiles actor operation | One bounded directory listing precedes parallel per-file work. Revisit this boundary if catalogs or remote volumes make enumeration measurably slow. |
focuspoints.json and settings reads | Detached utility task | Whole-file Data(contentsOf:) can block even for normally small JSON files. |
| ARW/NEF/DNG TIFF, MakerNote, and embedded-JPEG parsing | Runs within RAW-loader, extractor, scan, or package-test worker context | FileHandle and mapped/full-file fallback reads are synchronous. They must not be called directly from the main actor; full-file fallbacks make duration input-dependent. |
| Sorting, filtering, histogram math, and small result transforms | @concurrent | CPU work is bounded, does not wait on blocking APIs, and benefits from leaving the caller’s actor without requiring a dedicated blocking thread. |
| rsync startup | Synchronous argument/include-list preparation and executeProcess() on ExecuteCopyFiles’ main-actor method; output and completion are streamed asynchronously by RsyncProcessStreaming | executeProcess() launches and returns; it does not synchronously wait for rsync to finish. Include-list size and launch latency must stay bounded or be moved off the main actor. |
App-owned artifact and culling persistence actors also perform atomic reads/writes and directory maintenance. Serialization protects their state, but it does not make filesystem APIs suspend. Keep batches bounded and move any measured long operation to detached I/O while passing only Sendable values back to the owner.
@concurrent, Detached Tasks, And The GCD Bridge
These mechanisms are not interchangeable:
| Mechanism | Use it for | Do not use it as |
|---|
@concurrent | Bounded CPU work such as sorting, filtering, small transformations, or a short diagnostic calculation that should not inherit actor isolation | A general wrapper for ImageIO, full-file reads, RAW rendering, JPEG encoding, or other calls that may block for an unbounded time |
Task.detached | A contained filesystem or rendering operation where the caller passes immutable/Sendable inputs and awaits the result | A way to escape ownership, priority, cancellation, or Sendable rules; cancellation must still be checked and structured lifetime retained by awaiting .value where required |
CancellableImageIOWork | Reusable, cancellation-aware package APIs around nonsuspending ImageIO/Core Image work | Proof that the underlying call itself is cancellable; it only controls the waiter and checkpoints around the call |
| Dedicated process API | A long-running external command whose output, cancellation, and termination have their own lifecycle | Work to wait for synchronously on an actor or Swift task thread |
Quality Of Service
The chosen GCD QoS communicates user impact:
| Work | Typical QoS | Reason |
|---|
| On-demand thumbnails | userInitiated | User is scrolling or selecting images |
| Bulk cache warming | utility/background | Useful but should yield to direct UI work |
| JPEG export/cache warming | utility | Batch work that can run behind UI interaction |
| Memory diagnostics sampling | utility/detached | Should not block UI rendering |
The exact QoS is set in the package extractor or caller. When adding a new path, choose based on whether the user is waiting for the result right now.
Image Analysis
Sharpness scoring is adapted through RawCullPhotoAnalysisAdapter. Embedded-preview or demosaiced-RAW preparation runs away from the main actor, and the pipeline checks cancellation between decode, Vision, and scoring phases. Deep AI Review likewise marks decode/inference helpers @concurrent; any synchronous CIRAWFilter/CIContext portion must stay concurrency-limited because @concurrent alone does not turn rendering into a suspending operation.
The scoring image can come from:
| Source | Meaning |
|---|
embeddedPreview | Prefer embedded JPEG or ImageIO thumbnail for speed |
rawDemosaic | Use CIRAWFilter for a slower but more precise demosaiced thumbnail |
The code normalizes decoded images to 8-bit sRGB RGBA before the focus pipeline. That makes scoring less sensitive to source color space or bit depth and provides a clear Sendable/value boundary where possible.
Direct Binary Parsing
Sony, Nikon, and DNG parsers use FileHandle to read bounded leading regions
first, but some fallbacks read the full file to find later TIFF/MakerNote
structures or JPEG ranges. They locate focus-point data and embedded JPEG
offsets, then may read the selected JPEG byte range directly. These synchronous
operations must remain inside scan/parser/extractor worker contexts.
The parsers are written as stateless enums, so they do not need actor isolation.
Safe Rules For New Blocking Work
- Keep blocking APIs out of SwiftUI view bodies and
@MainActor methods. The narrow rsync launch path is an explicit exception only while launch remains short and non-waiting. - Put reusable blocking ImageIO/Core Image work in
RawParserKit behind an async API and the GCD continuation bridge. - Use detached I/O for isolated app-owned filesystem work, capture only Sendable values, and await the result when subsequent ownership or security-scope cleanup depends on completion.
- Use
@concurrent for bounded CPU work, not merely because a function is synchronous. - Bound catalog-wide decode/render fan-out with a limiter; executor choice does not impose backpressure.
- Add cancellation checkpoints before and after expensive framework calls.
- Convert non-Sendable image objects before crossing actor/task boundaries when needed.
- Put pure parsing or calculation logic in package code with fixtures and tests.
When A Synchronous Call Is Acceptable
A synchronous call is safe inside an actor only when all of these are true:
- it is small and bounded,
- it does not perform network/removable-volume I/O or decode/render a full image,
- it cannot expand from a small header/record into an unbounded full-file or directory operation,
- it does not capture mutable UI state,
- cancellation delay would not be visible to the user,
- multiplying it by the maximum actor/task-group concurrency still leaves cooperative threads available.
Examples include formatting values, cache-key construction, parsing a fixed-size in-memory record, or calculating display data. Move the work to detached I/O or CancellableImageIOWork when duration depends on file size, decoder behavior, volume latency, catalog size, or external-process completion. When uncertain, measure with a representative large RAW file and catalog; an actor is an ownership boundary, not a blocking-work queue.
Review Checklist
- Search for
CGImageSource, CGImageDestination, CIRAWFilter, CIContext, Data(contentsOf:), Data.write, FileHandle, directory enumeration, and process launch when auditing a new release. - Verify package extractors still go through
CancellableImageIOWork and catalog callers still apply decode limits. - Verify detached closures capture
Data, URL, scalar configuration, or other Sendable values rather than actor-owned CGImage/NSImage state. - Verify cancellation and completion races resume continuations exactly once;
CancellableImageIOWorkTests.swift covers this bridge. - Verify export/security-scope owners await detached writes before stopping access.
- Remove deprecated compatibility names from call sites and documentation as migrations complete.
14 - Documentation Update Plan
Prioritized backlog for keeping RawCull technical documentation aligned with the code
Documentation Update Plan
This page is the working backlog for future TechDocRawCull updates. It is
ordered by the risk that stale documentation will teach the wrong architecture,
not simply by the age or length of an article.
The intended reader understands Swift and SwiftUI at an intermediate level. Each
article should therefore explain ownership, data flow, cancellation,
persistence, and extension points before presenting low-level formulas or
implementation details.
Recently Completed Baseline
The following pages were reconciled with the RawCull source on 15 September 2026
and form the current learning path:
| Page | Current baseline |
|---|
| RawCull Tech Documentation | Repository map, composition root, architecture, and reading order |
| Burst Groups | Backend-selectable similarity artifacts, per-file persistence, cache schema 9, and the current workspace |
| Thumbnails and Scan Pipeline | Preload gating, request coalescing, replacement-safe identity, and current cache admission rules |
| Cache System | Representation-aware thumbnail caches and two-level similarity persistence |
| File Read and Write | Settings JSON, saved-data recovery, exports, security scopes, diagnostics, and rsync cleanup |
These pages still need review whenever their source areas change, but they are
not part of the immediate stale-document backlog below.
Priority Definitions
| Priority | Meaning | Target |
|---|
| P0 | The article may currently teach an incorrect runtime model or important invariant | Update before using it as an implementation guide |
| P1 | The article is broadly useful but lacks current ownership, UI flow, testing, or failure behavior | Update after P0 |
| P2 | The article is specialized or operational and should be checked against current packages, release tooling, or evidence | Update after core architecture pages |
| P3 | The content is stable process guidance with low architectural risk | Review when the workflow changes |
P0 — Correct The Core Runtime Model
Status: Completed
1. Concurrency
Page: Concurrency
Why first:
- The introduction still names the
RawCullAIModels branch rather than
documenting the current repository state. - The article predates the latest thumbnail contention work and should
explicitly include
ThumbnailPreloadGate, exact-key request coalescing,
waiter cancellation, and replacement-safe cache identity. - It should connect application termination, persistence flushing, catalog
security scope, JPG export scope, and rsync operation scope to their actual
owners.
Required update:
- Start at
RawCullApp, RawCullMainView, and the
@MainActor RawCullViewModel composition and presentation boundaries. - Add a hop diagram for catalog load, visible thumbnail demand, burst indexing,
and application termination.
- Distinguish actor serialization, bounded task groups, explicit
Task.detached, and framework callbacks. - Document generation checks, latest-wins behavior, continuation ownership, and
cancellation cleanup.
- Add a source-to-test table for concurrency invariants.
Completion evidence:
- Every named actor and task owner exists in the current source.
ThumbnailProviderTests, RawCullVerifyTestsConcurrencyTests,
RawCullVerifyTestsDataRaceDetectionTests, persistence tests, and
security-scope tests support the documented rules.
2. Focus Mask And Sharpness Overview
Page: Focus Mask and Sharpness
Why now:
- It is the bridge between the UI,
SharpnessScoringModel, FocusMaskModel,
PhotoAnalysisKit, saved culling data, and burst ranking. - Recent scoring settings, source selection, calibration, and cache-signature
changes should be reflected before readers use the detailed algorithm pages.
Required update:
- Add an ownership diagram from
SharpnessControlsView and scoring sheets
through the main-actor models into PhotoAnalysisKit. - Explain
SharpnessAnalysisDescriptor, effective thumbnail size, source
choice, calibration lifetime, and persistence validation. - Separate scalar sharpness, saliency evidence, focus-point evidence, and the
rendered focus mask.
- Document cancellation, bounded scoring, progress publication, and stale-result
prevention.
- Replace broad file lists with a guided “read these files in order” section.
3. Detailed Sharpness Scoring
Page: Detailed Sharpness Scoring
Required update:
- Revalidate every constant, default, formula, quality preset, source choice,
and score range against the pinned PhotoAnalysisKit revision.
- Label package-owned behavior separately from RawCull-owned orchestration and
UI normalization.
- Add one compact worked example for a medium-level reader before the
formula-by-formula reference.
- Link each major step to the test that protects it.
- Remove duplicated explanation already covered by the overview and retain this
page as the algorithm-level reference.
4. Detailed Focus Mask Computation
Page: Detailed Focus Mask Computation
Required update:
- Revalidate mask stages, region-selection rules, AF weighting, patch ranking,
thresholds, and debug modes against PhotoAnalysisKit.
- Explain which values change the scalar score, which change only mask
presentation, and which are calibration output.
- Add a data-shape diagram showing
CGImage/CIImage, analysis values, mask
output, and the SwiftUI overlay boundary. - Reduce repetition with the overview while preserving the step-by-step source
walkthrough.
P1 — Complete The Architecture Learning Path
Status: Completed
5. AI Section Overview
Page: Artificial Intelligence in RawCull
Required update:
- Present
RawCullAIIntegration as the composition root and list the narrow
services passed into feature models. - Separate burst similarity, semantic search, and Deep Review; they use related
packages but have different capability and persistence rules.
- Explain Vision availability, optional CLIP selection, CLIP-to-Vision recovery,
segmentation model selection, and capability refresh.
- Align its learning order with the main documentation index and remove
duplicated model-download instructions.
6. CLIP Runtime Integration
Page: How RawCull Loads and Uses CLIP
Required update:
- Verify startup behavior, managed model locations, bundle validation, provider
reuse, model fingerprints, and settings callbacks.
- Add the current boundary between burst similarity artifacts and
semantic-search artifacts.
- Document partial CLIP generation, whole-batch Vision fallback, diagnostic
logging, and descriptor validation.
- Explain per-file artifact hydration and why changing backend descriptors
invalidates reuse.
7. Package Overview And Reading Order
Page: RawCull Packages
Required update:
- Derive the package list and revisions from
Package.resolved. - Show which products are imported by the app and which types form each
boundary.
- Add a dependency-direction diagram covering RawCullCore, RawParserKit,
PhotoAnalysisKit, PhotoAIKit, and the rsync support packages.
- State that package repositories are separately versioned and are not source
snapshots inside TechDocRawCull.
8. PhotoAIKit
Page: How PhotoAIKit Is Constructed
Required update:
- Compare the article with the exact pinned package revision.
- Recheck product names, contract types, backend actors, artifact descriptors,
storage APIs, fallback behavior, and segmentation workflows.
- Add a RawCull integration section mapping package protocols to
RawCullAIIntegration, SimilarityScoringModel, semantic search, and Deep
Review. - Identify which behavior belongs to the package and which policy remains in
RawCull.
9. PhotoAnalysisKit
Page: How PhotoAnalysisKit Is Constructed
Required update:
- Reconcile the package facade, analysis descriptors, presets, batch limits,
calibration, focus evidence, and mask APIs with the pinned revision.
- Add call maps from RawCull’s sharpness and focus models into the package.
- Link package tests for scalar scoring, cancellation, configuration identity,
and mask rendering.
10. RawCullCore
Page: How RawCullCore Is Constructed
Required update:
- Verify domain models,
FileItem typealias boundaries, burst grouping/ranking
defaults, review states, and histogram behavior. - Explain why pure
nonisolated value logic belongs here while orchestration
and persistence remain in the app. - Add extension guidance for new grouping evidence and cache-version
consequences.
11. RawParserKit
Page: How RawParserKit Is Constructed
Required update:
- Verify format registration, metadata normalization, thumbnail/preview
strategies, coalescing, decode limits, and cancellation bridges.
- Map
RawParserKitImageLoader to the package facade and show where
non-Sendable images are consumed or converted. - Include the current Sony and Nikon behavior without implying that app code
should call vendor conformers directly.
12. Sony/Nikon MakerNote Parser
Page: Sony/Nikon MakerNote Parser
Required update:
- Reconcile parser type names and file locations with RawParserKit.
- Walk one Sony and one Nikon focus-location result through normalization into
FileItem and focus UI. - Document fallback to EXIF subject area and catalog-wide
focuspoints.json
behavior. - Add a checklist and tests required when introducing another RAW format.
P1 — Operational Correctness References
Status: Completed
13. Security-Scoped URLs
Page: Security-Scoped URLs
Required update:
- Keep the already-current AI indexing and semantic-search scope explanation.
- Add the selected JPG export destination lifetime and app-termination
persistence flush.
- Cross-check rsync bookmark fallback, idempotent cleanup, and exact ownership
of every successful scope start.
- Add a table for catalog, scan, export, copy, diagnostics, and AI operations
showing owner, start, stop, and failure cleanup.
14. Memory Pressure
Page: Memory Pressure
Required update:
- Recheck adaptive cache recommendations, user maxima, warning/critical
responses, and recovery behavior.
- Connect pressure state to both grid and preview caches and to diagnostics
counters.
- Explain the lock-backed synchronous read without teaching that every cache
operation bypasses actor isolation.
- Add the tests and diagnostic measurements used to validate limit changes.
15. Synchronous Code
Page: Synchronous Code
Required update:
- Audit all current blocking ImageIO, filesystem, RAW parsing, JPEG encoding,
and process operations.
- Distinguish short
@concurrent work from operations intentionally moved to a
detached task or dedicated GCD continuation bridge. - Add decision rules for when a synchronous call is safe inside an actor and
when it would occupy Swift’s cooperative executor too long.
- Remove types or paths that no longer exist.
P2 — AI Distribution, Evidence, And Release Procedures
Status: Completed
16. AI Model Downloads
Page: AI Model Download Service
Required update:
- Reconcile the procedure with the current download catalog, downloader target,
managed locations, activation callbacks, and release metadata tests.
- Separate runtime architecture from release-operator commands.
- Add failure/retry behavior and the user-visible capability states.
17. Publishing New AI Models
Page: Publishing New RawCull AI Models
Required update:
- Verify archive names, manifest schema, checksums, release tags, model
identities, and staging paths against
ModelAssets and current release tests. - Replace any historical one-off commands with parameterized examples or clearly
label them as records.
- Add a final reproducibility and licence gate before publishing.
18. AI Licence And Provenance Procedure
Page: AI Model Licence and Provenance Clearance
Required update:
- Separate current legal/provenance status from the reusable clearance
procedure.
- Verify notices, provenance JSON, upstream licences, acceptance requirements,
and distribution restrictions for every shipped model.
- Add an evidence date and owner to decisions that can expire or change.
- Keep legal conclusions explicitly evidence-based and avoid inferring
permission from model availability.
19. Evaluating CLIP Models
Planned page: Evaluating CLIP Models (not yet present)
Required update:
- Verify package revisions, fixture identities, scripts, commands, thresholds,
and report paths.
- Separate parity testing, semantic retrieval evaluation, performance
measurement, and RawCull integration testing.
- Add a reproducibility checklist including hardware, OS, toolchain, model
fingerprint, and immutable fixture digest.
20. CLIP Evaluation Results
Planned page: CLIP Model Evaluation Results (not yet present)
Required update:
- Treat this as a dated evidence report rather than timeless architecture.
- Record exact inputs, model fingerprints, query set, metrics, hardware, and
report generation date.
- Link conclusions to generated artifacts and clearly distinguish measured
results from recommendations.
- Add a superseded-results policy so later evaluations do not silently overwrite
historical evidence.
P3 — Stable Workflow Guidance
21. Repository Git Workflow
Page: Repository Git Workflow
Required update:
- Confirm that the documented rebase and fast-forward policy still matches
repository practice.
- Add the documentation validation commands: Prettier, Hugo build, and
internal-link check.
- Remove duplicated Git basics if the page is intended only for this
repository’s policy.
- Review when branch protection, deployment, or contribution rules change.
Proposed New Pages
These pages should be added only after the existing P0 and P1 articles are
accurate.
| Priority | Proposed page | Purpose |
|---|
| P1 | swiftuiarchitecture.md | Main window modes, NavigationSplitView composition, environment injection, sheets, overlays, commands, and reusable inspection views |
| P1 | persistence.md | CullingModel, saved-file schema, backup/corruption recovery, debounced writes, flush-on-termination, and migration rules |
| P2 | testing.md | Test plans, smoke/performance manifests, package tests, isolation helpers, fixtures, and how tests encode architecture invariants |
| P2 | diagnostics.md | Memory, similarity, contention, and RAW diagnostics; log locations, privacy boundaries, and troubleshooting workflow |
Standard Required For Every Update
Every revised architecture article should contain:
- Purpose and boundary — what the subsystem owns and deliberately does not
own.
- Source map — exact repository-relative files and package revision where
relevant.
- Read order — the shortest path through the code for a new contributor.
- End-to-end flow — trigger, main-actor orchestration, background owner,
persistence, and UI publication.
- State and lifetime — actor ownership, cancellation, generation guards,
security scopes, and cleanup.
- Failure behavior — what the user sees and what remains recoverable.
- Cache or persistence identity — descriptors, fingerprints, versions, and
invalidation rules.
- Tests — executable evidence for important invariants.
- Change checklist — related files, versions, docs, and tests to update
together.
- Last reviewed date — the date the article was checked against source,
not merely reformatted.
Avoid absolute developer-machine paths, undocumented source snapshots,
unverified constants, and claims that await automatically moves work to a
background thread.
Review Triggers
Update the relevant article in the same change whenever any of these occur:
- a source file is renamed or ownership moves between view, model, actor, and
package;
- a package revision changes a public contract or default;
- a cache key, schema, descriptor, algorithm version, or persistence format
changes;
- a new task, actor, continuation, security scope, or cancellation path is
introduced;
- a new model backend, RAW format, view mode, export mode, or diagnostics store
is added;
- tests establish a new invariant that the current article does not explain.
After each documentation batch, run the configured formatter, render the
complete Hugo site, validate internal links, and check the diff for stale
filenames and obsolete constants.
15 - Sony, Nikon, and DNG Metadata Parsers
RawParserKit 1.3.0, revision
d2175ed880d39021bdb5f5a2a842b460af0b316c, provides one neutral result
shape for Sony ARW, Nikon NEF, and Adobe DNG autofocus metadata. RawCull enters
through RawImageLoader.metadata(for:) or RawFormatRegistry; diagnostics and
package tests may call a format parser directly.
Current Source Map
| Area | RawParserKit file |
|---|
| Neutral format contract and registration | Sources/RawParserKit/RawFormat.swift, RawFormatRegistry.swift |
| Normalized focus value and metadata snapshot | Sources/RawParserKit/BrowserFocusPoint.swift, BrowserExifInfo.swift |
| Facade and EXIF fallback | Sources/RawParserKit/RawImageLoader.swift |
| Sony | SonyMakerNoteParser.swift, SonyRawFormat.swift, SonyThumbnailExtractor.swift, SonyEmbeddedJPEGExtractor |
| Nikon | NikonMakerNoteParser.swift, NikonRawFormat.swift, NikonThumbnailExtractor.swift, NikonEmbeddedJPEGExtractor |
| DNG | DNGMakerNoteParser.swift, DNGRawFormat.swift, DNGThumbnailExtractor.swift, DNEmbeddedJPEGExtractor.swift |
| RawCull consumer | RawCull/Model/RawImageLoading.swift, RawCull/Actors/ScanFiles.swift |
RawFormatRegistry.all registers SonyRawFormat for .arw, NikonRawFormat
for .nef, and DNGRawFormat for .dng. Each conformer supplies focus-point
parsing, thumbnail and embedded-preview extraction, compression labels, and RAW
size thresholds.
Shared Focus Contract
All three parser paths return:
imageWidth imageHeight focusX focusY
RawFocusPoint validates four numbers, positive dimensions, and normalized
coordinates in 0...1. The RawCull adapter converts that value to CGPoint for
FileItem.afFocusNormalized while retaining the four-number compatibility
string for the focus overlay.
flowchart LR
Parser["ARW / NEF / DNG parser"] --> Format["RawFormat.focusLocation"]
Format --> Meta["RawImageLoader metadata + RawFocusPoint"]
Meta --> Adapter["RawParserKitImageLoader"]
Adapter --> Item["FileItem.afFocusNormalized"]
Adapter --> Overlay["FocusPointsModel and overlay"]
Item --> Analysis["focus evidence and ranking"]For example, "6000 4000 3000 2000" normalizes to (0.5, 0.5). A Nikon
AFInfo2 value such as "8256 5504 2064 1376" normalizes to (0.25, 0.25).
The examples describe the contract; package tests use synthetic TIFF/MakerNote
structures rather than camera files.
Fallback Order
RawImageLoader.metadata(for:) asks the registered format for a focus location
first. If none is produced, it reads kCGImagePropertyExifSubjectArea, treats
the first two numbers as pixel X/Y, and normalizes them against image width and
height. Invalid dimensions or out-of-range values produce no point.
RawCull then has a catalog-wide compatibility fallback: ScanFiles reads
focuspoints.json only when the entire native-point collection is empty. If
even one file has a native MakerNote or EXIF point, JSON is not merged into the
partly populated result. JSON supplies overlay compatibility data; it does not
retroactively populate every FileItem.afFocusNormalized.
- Sony follows TIFF IFD0 to EXIF and the Sony MakerNote IFD, including
FocusLocation
0x2027, and reports embedded JPEG candidates. - Nikon validates the Type-3 MakerNote and reads supported AFInfo2
0x00B7
layouts. Unsupported layouts return nil. - DNG walks TIFF IFD0 and SubIFDs. Standards-classified files use
NewSubFileType plus Compression to distinguish thumbnails/previews from
JPEG-compressed raw strips. Files without NewSubFileType retain a positional
fallback. DNG compression labels include uncompressed, JPEG, Deflate,
PackBits, Lossy DNG, and JPEG XL.
The DNG container is camera-neutral, so its size classes use generic megapixel
thresholds with overrides for known high-resolution, full-frame, and smaller
sensor camera families.
Tests At The Pinned Revision
| Test | Contract |
|---|
SonyMakerNoteParserTests.swift | Sony focus tags, offsets, malformed input, and embedded JPEG discovery |
NikonMakerNoteParserTests.swift | Nikon Type-3/AFInfo2 layouts, byte order, and embedded JPEG discovery |
DNGMakerNoteParserTests.swift | DNG TIFF/SubIFD focus and preview discovery, including malformed data |
DNGRawFormatTests.swift | DNG compression labels, size classes, and format behavior |
RawFormatRegistryTests.swift | Case-insensitive ARW/NEF/DNG dispatch and unregistered formats |
- Add a stateless
RawFormat conformer with extensions, display name,
extraction, focus location, compression labels, and size thresholds. - Register it in
RawFormatRegistry.all; keep app code on the registry and
RawImageLoader facade. - Normalize autofocus output to
"width height x y" and prove bounds,
endianness, truncation, missing tags, and unsupported versions are safe. - Add synthetic parser, registry, metadata-fallback, preview, orientation,
cancellation, and diagnostics tests.
- Verify the RawCull adapter maps the result into the same visual coordinate
system and preserves the catalog-wide JSON fallback rule.
16 - Repository Git Workflow
Repository Git Workflow: Linear History
This page documents the preferred repository workflow for RawCull documentation changes: branch-based development, signed commits when configured, rebasing onto main, and fast-forward integration without merge commits.
Create a Branch
git checkout -b <new-branch>
git push --set-upstream origin <new-branch>
Daily Workflow
You always work on a dedicated branch. Commit often as you make progress.
Stage changes
Commit
If you use the Claude CLI, it can generate a Conventional Commits message from the staged diff:
git commit -m "$(git diff --staged | claude -p 'Write a short conventional commit message. Output only the message, nothing else.')"
Claude pipes the staged diff into the claude CLI and returns a single-line message such as feat(cache): add LRU eviction for thumbnail layer. The -p flag runs Claude non-interactively (print mode) so the output can be captured directly into -m.
If you want to review the message before committing, capture it first:
MSG=$(git diff --staged | claude -p 'Write a short conventional commit message. Output only the message, nothing else.')
echo "$MSG"
git commit -m "$MSG"
Repeat staging and committing as often as needed while working.
Push
git push origin <your-branch>
Keep your branch current
If others have pushed to main while you were working, rebase your commits on top of their work rather than merging.
git fetch origin
git rebase origin/main
This puts your commits aside, fast-forwards your branch to the latest main, then replays your commits on top — keeping history a straight line.
Integrate into main (fast-forward)
When your branch is finished and ready to ship, follow this procedure to keep history linear.
1. Update main from the server
git checkout main
git pull --rebase origin main
2. Rebase your branch onto the fresh main
git checkout <your-branch>
git rebase main
3. Fast-forward merge into main
--ff-only makes Git abort instead of creating a merge commit.
git checkout main
git merge --ff-only <your-branch>
4. Push main to GitHub
5. Delete the branch locally
git branch -d <your-branch>
6. Delete the branch on GitHub
git push origin --delete <your-branch>
Using git fetch and git diff
The git fetch command is used to update your local repository with the latest changes from the remote repository without merging them. You can then use git diff to compare the branches.
Step 1: Fetch the Latest Changes
Fetch the latest changes from the remote repository to ensure you have the most up-to-date information.
Step 2: Compare the Branches
Use the git diff command to compare your local branch with the remote branch.
git diff <local-branch> origin/<remote-branch>
For example, if you want to compare your local main branch with the remote main branch:
git diff main origin/main
Using git log
The git log command can be used to compare commit histories between your local and remote branches. This is useful for seeing which commits are present in one branch but not the other.
Fetch the Latest Changes:
Ensure your local repository is updated with the latest changes from the remote repository.
Compare Commit Histories:
Use the git log to see the differences in commit histories.
git log <local-branch>..origin/<remote-branch>
For example, to compare your local main branch with the remote main branch:
git log main..origin/main
You can also reverse the comparison to see commits in the remote branch that are not in the local branch:
git log origin/main..main
Using git status
The git status command provides a quick summary of the differences between your local branch and the remote branch.
Fetch the Latest Changes:
Update your local repository with the latest changes from the remote repository.
Check the Status:
Use git status to see the differences between your local branch and the remote branch.
The output will show messages like “Your branch is ahead of ‘origin/’ by X commits” or “Your branch is behind ‘origin/’ by X commits”, indicating the differences.
Pull and track a remote repository
The error means your local branch has no upstream tracking set. Fix it with:
git branch --set-upstream-to=origin/<your-branch> <your-branch>
git pull
Or do both in one step:
git pull origin <your-branch>
To avoid this in the future, whenever you create or checkout a new local branch that should track a remote, use:
git checkout --track origin/<your-branch>
Verify linear history
git log --oneline --graph --decorate
A clean linear history shows a straight vertical line with no merge nodes.
Re-sign the last commit without changing its message
git commit --amend --no-edit -S
Push after re-signing (the commit hash changed)
git push origin <your-branch> --force-with-lease
Fix GPG if signing fails
git config --global gpg.format openpgp
git config --global gpg.program gpg
# Confirm the Key ID is correct (no leading '0x')
git config --global user.signingkey <YOUR_KEY_ID>
# Restart the agent
gpgconf --kill gpg-agent
One-time setup
Run these commands once per machine to configure Git correctly for this workflow.
1. Set your identity
git config --global user.name "Your Name"
git config --global user.email "you@example.com"
2. Verify the remote uses SSH
If the URL starts with https://, switch it to SSH:
git remote set-url origin git@github.com:<user>/<repo>.git
# Find your GPG Key ID (the 16-character code after '/' on the 'sec' line)
gpg --list-secret-keys --keyid-format LONG
# Tell Git which key to use (omit the leading '0x')
git config --global user.signingkey <YOUR_KEY_ID>
# Auto-sign all commits and tags
git config --global commit.gpgsign true
git config --global tag.gpgSign true
git config --global gpg.program gpg
4. SSH connection
- Go to github.com/settings/keys and delete the old key.
- Copy your current public key to the clipboard:
pbcopy < ~/.ssh/id_ed25519.pub
- On GitHub: Settings → SSH and GPG keys → New SSH key → paste → Save.
- Test again:
5. Test the SSH connection to GitHub
A successful response looks like:
Hi <username>! You've successfully authenticated, but GitHub does not provide shell access.
If you get a Permission denied (publickey) error, the most likely cause is a stale key.
6. Enforce linear history (no merge commits)
# Always rebase instead of merge when pulling
git config --global pull.rebase true
# Refuse any merge that would create a merge commit
git config --global merge.ff only
# Simplify first push of a new branch (Git ≥ 2.37)
git config --global push.autoSetupRemote true
17 - Security-Scoped URLs
Security-Scoped URLs
RawCull is a sandboxed macOS app. Any access outside the app container must come from user consent, usually a file/folder picker. RawCull uses two security-scope patterns:
- active catalog access for browsing/culling and as the rsync source,
- a persistent bookmark for the rsync destination.
Source Map
| Area | Files |
|---|
| Active catalog scope | RawCullViewModel.swift, RawCullViewModel+Catalog.swift, RawCullApp.swift |
| Catalog scan scope | Actors/ScanFiles.swift |
| CLIP indexing | RawCullViewModel+Similarity.swift, SimilarityScoringModel.swift, RawCullVisionSimilarityService.swift |
| Semantic search | RawCullViewModel+Similarity.swift, SimilarityScoringModel.swift, RawCullSemanticSearchService.swift |
| Similarity artifact cache | Intelligence/Persistence/PerFileAnalysisArtifactStore.swift |
| Copy-folder bookmarks | Views/CopyFiles/OpencatalogView.swift, SourceAndDestinationSection.swift |
| rsync runtime scope | Model/ParametersRsync/ExecuteCopyFiles.swift |
| Selected JPG export | ExtractJPGsSheetView.swift, RawCullViewModel+Thumbnails.swift, ExtractAndSaveJPGs.swift, SaveJPGImage.swift |
| App termination | Main/RawCullApp.swift, CullingModel.swift |
API Basics
The core calls are:
let ok = url.startAccessingSecurityScopedResource()
url.stopAccessingSecurityScopedResource()
Persistent access is stored as bookmark data:
let data = try url.bookmarkData(options: .withSecurityScope, ...)
let url = try URL(resolvingBookmarkData: data, options: .withSecurityScope, ...)
Every successful startAccessing... must eventually be paired with stopAccessing....
Active Catalog Scope
The catalog browsing flow is owned by RawCullViewModel.
sequenceDiagram
participant UI as Sidebar picker
participant VM as RawCullViewModel
participant Work as Scan/thumbnail/export work
UI->>VM: startCatalogLoad(source)
VM->>VM: cancelCatalogLoad()
VM->>VM: startSecurityScopedAccess(url)
VM->>Work: scan and preload
Work-->>VM: results/progress
VM->>VM: stopActiveSecurityScopedAccess() on cancel/empty/deinit/app cleanupstartSecurityScopedAccess(for:) is idempotent for the currently active URL. If a different catalog is selected, it stops the previous active scope before starting the new one.
cancelCatalogLoad() releases the active scope and cancels related work. An empty scan also releases it. RawCullViewModel.deinit is the final defensive release.
Catalog changes are persistence boundaries: startCatalogLoad(for:) waits for CullingModel.flushPersistence() before it cancels the old catalog and its scope. If the flush fails, RawCull restores the previous selection and keeps the old catalog active.
ScanFiles Scope
ScanFiles.scanFiles(url:onProgress:) also starts and stops access around directory scanning:
let didStartSecurityScope = url.startAccessingSecurityScopedResource()
defer {
if didStartSecurityScope {
url.stopAccessingSecurityScopedResource()
}
}
This is a local defensive scope for the scan actor. It stops only when its own start succeeded. The broader catalog scope remains owned by RawCullViewModel so later preload, diagnostics, export, zoom, and AI work can still access files while the catalog is active.
Scope Ownership Matrix
The owner is the component that records a successful start and is therefore responsible for the matching stop. A borrower may use URLs covered by a longer-lived owner, but must not stop that owner’s scope.
| Operation | Scope owner | Start | Stop | Failure and cancellation cleanup |
|---|
| Active catalog | RawCullViewModel | startSecurityScopedAccess(for:) before catalog work | Catalog cancel/change, empty scan, successful app termination, or deinit | A failed start is not recorded. Switching first flushes culling persistence; a failed flush retains the old catalog and scope. |
| Directory scan | ScanFiles.scanFiles | Local startAccessingSecurityScopedResource() | defer, but only when the local start returned true | defer covers success, thrown filesystem errors, cancellation, and early return. The view model’s broader scope is not stopped. |
| Selected JPG export | RawCullViewModel.startSelectedJPGExtraction | Start the chosen destination immediately before creating ExtractAndSaveJPGs | On return from extractAndSavejpgs(), before publishing completion or failure UI | Failed destination start aborts without a stop. Per-file failures are collected; the operation-level stop still runs after the actor returns. Source reads borrow the active catalog scope. |
| rsync copy | ExecuteCopyFiles | Start the selected catalog URL, then resolve and start destBookmark | Idempotent cleanup() after normal completion, close/cancel, startup failure, launch failure, or deinit | There is no direct-path fallback. If destination setup fails, cleanup stops the already-started source. didCleanUp prevents duplicate stops and include-file removal. |
| AI indexing | Active catalog (RawCullViewModel) | No per-file start; indexing borrows the selected directory scope | No per-file stop | Index cancellation stops AI work, not the catalog scope. Catalog cancellation/change releases the owner scope after cancelling related work. |
| Semantic query | None for ranking; active catalog remains open for follow-on actions | No start; ranking reads hydrated in-memory artifacts | No stop | Query cancellation discards query work. Any subsequent preview/export uses the appropriate catalog or export scope. |
CLIP Indexing and Semantic Search
CLIP indexing and semantic search use the active catalog scope differently. Indexing reads source images, while a search query operates on cached embeddings.
flowchart LR
A["Active security-scoped catalog"] --> B["FileItem URLs"]
B --> C["Decode RAW thumbnail, max 512 px"]
C --> D["CLIP image encoder"]
D --> E["Validated similarity artifact"]
E --> F["Application Support cache"]
Q["Text query"] --> T["CLIP text encoder"]
F --> S["Cosine similarity ranking"]
T --> S
S --> R["Ranked catalog selection"]CLIP indexing requires the catalog scope
RawCullViewModel.indexSimilarity() first hydrates reusable artifacts and then asks SimilarityScoringModel.indexFiles(_:) to generate any missing or stale artifacts. Each FileItem becomes an AIImageSource containing the file URL.
For an artifact that must be generated, RawCullSimilarityImageDecoder reads the source URL. It first asks RawParserKitImageLoader for a thumbnail with a maximum dimension of 512 pixels and then tries ImageIO as a fallback. Because these URLs point into the user-selected catalog, decoding depends on the catalog directory’s security-scoped access still being active.
RawCull does not call startAccessingSecurityScopedResource() for every image. Access was already started for the selected directory by RawCullViewModel, and that scope covers its files. The view model deliberately keeps the directory scope open after the initial scan so indexing, thumbnail generation, previews, exports, and other catalog operations can read the same URLs.
The decoded image is passed to the selected local similarity backend. When the selected backend is CLIP, PhotoAIKit creates a normalized image embedding. Semantic-search coverage can only be populated when the active similarity backend produces artifacts compatible with the selected CLIP semantic-search backend. If Vision similarity is selected, the UI asks the user to enable Use selected CLIP model for similarity before building missing semantic-search artifacts.
Successfully validated artifacts are written one file at a time to:
~/Library/Application Support/RawCull/AnalysisArtifacts/Similarity/
In the sandbox, that resolves inside RawCull’s container. It is app-owned storage and does not need a security-scoped URL. Cache records are keyed and validated against the source fingerprint, artifact schema, model/backend descriptor, and RawCull’s embedding pipeline signature. A moved, renamed, changed, incompatible, or corrupt source is therefore treated as a cache miss and must be indexed again while the catalog scope is active.
Semantic search reuses cached CLIP artifacts
Opening a catalog hydrates compatible artifacts from the app-owned cache into semanticArtifacts. A semantic query then:
- applies the ordinary catalog admission rules, such as filename and rating filters,
- keeps only files with an artifact compatible with the currently selected CLIP backend,
- encodes the literal text query with the local CLIP text encoder,
- computes cosine similarity between that temporary text embedding and the cached image embeddings,
- sorts the matches and exposes the selected highest-ranked files as the active catalog working set.
searchSemantically(for:) and rankSemantically(query:files:) do not decode RAW files, generate image embeddings, or read image contents from the catalog. The text-query embedding exists only for that search call and is not persisted. Consequently, semantic ranking itself does not acquire a new security scope; it uses in-memory artifacts that were restored or created earlier.
The catalog scope nevertheless remains active during semantic search. Ranked files can immediately flow into preview, zoom, export, culling, burst, or Deep Review operations that do need their source URLs. Clearing a search changes the working set, not the security-scope owner or lifetime.
Scope lifetime for the AI workflow
sequenceDiagram
participant User
participant VM as RawCullViewModel
participant Index as CLIP indexing
participant Cache as App-owned artifact cache
participant Search as Semantic search
User->>VM: Select catalog
VM->>VM: startAccessing catalog URL
VM->>Cache: Hydrate compatible artifacts
User->>Index: Index Similarity
Index->>VM: Read source URLs under active scope
Index->>Cache: Persist validated image artifacts
User->>Search: Submit text query
Search->>Cache: Use hydrated CLIP artifacts
Note over Search: No source decoding and no new scope
User->>VM: Close, cancel, or select another catalog
VM->>VM: stopAccessing catalog URLCopy Workflow Bookmark
The copy workflow reuses the active catalog as its source and persists only the
destination. OpencatalogView creates destBookmark when the user picks that
folder.
flowchart TD
A["User picks destination"] --> B["startAccessing"]
B --> C["bookmarkData(.withSecurityScope)"]
C --> D["UserDefaults destBookmark"]
D --> E["stopAccessing"]
E --> F["Later: ExecuteCopyFiles resolves bookmark"]
F --> G["start destination scope during rsync"]
S["Selected catalog URL"] --> H["start source scope during rsync"]
G --> I["cleanup stops both scopes"]
H --> IIf the selected catalog scope cannot be started, the user is asked to reopen the
catalog. If destBookmark is missing or cannot be resolved, the user must
reselect the destination. A stale bookmark is regenerated while its resolved
scope is active. If source access succeeds but destination access fails,
cleanup() releases the source before returning the startup error.
rsync Runtime Cleanup
ExecuteCopyFiles stores the accessed URLs in:
sourceAccessedURL,destAccessedURL.
cleanup() finishes the progress stream, stops both security-scoped resources, clears process references, and is guarded by didCleanUp so multiple termination paths are safe.
close() sets isClosing, cancels the process, and calls cleanup. Normal
termination constructs a CopyDataResult containing the output, typed outcome,
and immutable CopyOperation source/destination snapshot, invokes completion,
and then cleans up. There is no timing-delay dependency in the completion path.
The include list is written under Application Support/RawCull/CopyLists, not the user-selected source or destination. Cleanup removes the per-operation list on every path after it has been created.
Selected JPG Export Destination
The export sheet can use an existing catalog or a folder returned by its Choose… file importer. extractJPGDestination remembers that choice for the lifetime of the RawCullViewModel; it is not a persistent rsync-style bookmark.
Pressing Extract starts a separate security scope for the destination immediately before ExtractAndSaveJPGs begins. The scope remains active for the entire batch, including detached atomic writes performed by SaveJPGImage, and is stopped when extractAndSavejpgs() returns. The active source catalog scope remains a different ownership unit and covers RAW/JPEG reads. If source and destination happen to be the same URL, each successful start still belongs to its own operation and must receive its own stop.
Destination access failure prevents actor creation and presents Export Not Started. Individual extraction or write failures do not shorten the destination lifetime: the actor returns an aggregate result, the view model stops the destination scope, and then it presents Export Incomplete if needed.
App Termination And Persistence
AppDelegate.applicationShouldTerminate(_:) returns .terminateLater and starts one termination task. That task awaits cullingModel.flushPersistence() before releasing the active catalog scope. A second termination request while the task is running also returns .terminateLater rather than starting another flush.
If persistence succeeds, the app stops the active catalog scope and replies
true to AppKit. If persistence fails, a modal recovery choice offers retry,
cancel quitting, or Quit Without Saving. Retry repeats the flush; cancel
keeps the app and scope alive; discard releases the scope and terminates while
explicitly acknowledging unsaved culling changes. Only one termination task is
active. Export and rsync retain their own cleanup ownership.
File Writes
Writes outside the app container require an active user-granted scope. The main examples are:
| Write | Scope source |
|---|
| Extracted JPEG sidecars next to RAW files | active catalog scope |
| Selected JPG export folder | operation-owned destination scope |
| rsync destination writes | destBookmark scope |
| rsync include file | app Application Support folder, no external scope required |
App-owned JSON/cache files under Application Support or Caches do not need security-scoped access.
What To Check When Changing This Area
- Keep one clear owner for each long-lived scope.
- Pair every successful start with a stop on all exit paths.
- Keep catalog switching and app termination behind a successful culling-state persistence flush.
- Keep the selected JPG destination scope open until the whole export actor returns; do not stop it after scheduling detached writes.
- Keep the active catalog scope alive for CLIP indexing and for operations launched from semantic-search results.
- Do not add per-file security-scope calls inside the CLIP indexer; the selected catalog directory owns that access.
- Keep semantic ranking cache-only. If it begins decoding source images, its security assumptions and UI behavior must be revisited.
- Store similarity artifacts and query-independent embeddings in app-owned storage, not beside the RAW files.
- Use bookmarks for persistent copy-folder access, not for every temporary catalog scan.
- When adding a new file write, ask whether it targets the app container or a user folder.
- If copy closes early, verify
ExecuteCopyFiles.cleanup() still runs exactly once. - Test partial rsync startup (source succeeds, destination fails) and partial JPG export failures as ownership cases, not only as UI errors.