This is the multi-page printable view of this section. Click here to print.

Return to the regular view of this page.

Artificial Intelligence

RawCull AI architecture, model downloads, and Objects test-release status.

Artificial Intelligence in RawCull

This section explains how RawCull uses reusable AI components without allowing model-runtime details to spread through the application. It is written as a learning path: first understand the boundary between RawCull and PhotoAIKit, then study the package itself, and finally follow CLIP from application startup to a persisted similarity artifact.

The source for this section comes from the pinned PhotoAIKit dependency and the RawCull repository. Paths beginning with Sources/ refer to PhotoAIKit at the revision in Package.resolved. RawCull intelligence code lives under RawCull/Intelligence; app composition and presentation remain under RawCull/Main, RawCull/Model, and RawCull/Views.

What This Section Documents

DocumentMain questionStart here when
This overviewWhere does AI belong in the system?You need the vocabulary and responsibility split
The RawCull AI RuntimeHow are providers and long-lived features assembled?You are tracing startup, refresh, or service replacement
AI Models in RawCullHow do CLIP, Vision, SAM 3, Qwen, and Objects work in the app?You are tracing an analysis from input to result
Download AI modelsHow are the three release packs rebuilt?You are preparing source weights, conversions, or archives
AI Model Licence and Provenance ClearanceWhat evidence is required before a model can ship?You are reviewing licences, provenance, or release readiness
Publishing and Testing RawCull AI ModelsHow are packs published and Objects tested?You are preparing a model release or updating the download manifest

PhotoAIKit and RawCull use these AI backends:

  • CLIP image embeddings for visual similarity and semantic search.
  • SAM 3 and EfficientSAM subject segmentation for subject masks.
  • Apple Vision feature prints for always-available image similarity.
  • Qwen3-VL for local photo analysis and for concept discovery and assessment of numbered SAM 3 objects.

The 3.2.6 production download catalog exposes DataComp CLIP, Meta SAM 3, and Qwen3-VL-2B-Instruct. OpenAI CLIP remains excluded and EfficientSAM is not a production download. SAM 3 requires acceptance of its verified bundled licence before download. The App Store build uses Apple-hosted Managed Background Assets; the Direct/Developer ID build retains a self-hosted v3 manifest. See AI Model Downloads for the exact distinction.

Vision is the startup and service-selection fallback: RawCull uses it when CLIP is disabled or the selected CLIP bundle cannot produce a validated provider. A selected CLIP indexing pass keeps its valid per-file artifacts and records the files that fail; it does not mix Vision artifacts into that pass or automatically rerun the whole batch.

The AI Models in RawCull guide follows code connected to similarity, semantic search, burst analysis, Deep Review, Qwen, and Objects. The PhotoAIKit architecture guide covers SAM 3 contracts, workflows, and storage so that the package design is understandable. Not every reusable package capability is necessarily exposed as a finished RawCull user workflow.

Objects test-release status

The maintainer is preparing a user test of AI Analysis → Objects. The feature combines Qwen concept discovery, separate SAM 3 masks, and Qwen assessment of numbered crops. The structural decoder checks board IDs and response fields; it cannot verify that generated text matches the photograph. A recent two-puffin result had two retained objects but an invented third bird in its summary, and a separate two-puffin result has an unresolved crop/description mismatch. Treat Objects output as advisory and verify it against the source. See Publishing and Testing RawCull AI Models for the known issues, tester checks, and release evidence still needed.

The Central Design Idea

RawCull and PhotoAIKit answer different kinds of questions.

PhotoAIKit asks:

  • What does an image source look like at a package boundary?
  • How is a model bundle validated and identified?
  • How does a backend produce and compare a similarity artifact?
  • How is bounded indexing or segmentation orchestrated?
  • How can reusable artifacts and masks be encoded or cached?

RawCull asks:

  • Where are models installed for this application?
  • How is a Sony, Nikon, or DNG RAW file decoded for AI input?
  • Which backend did the user request?
  • When should a catalog be indexed or reindexed?
  • How do similarity distances affect burst grouping and culling?
  • What state and wording should SwiftUI present?

This is dependency inversion in practical form. PhotoAIKit defines small protocols such as ImageDecoding, ImageSimilarityArtifactProviding, and ImageSimilarityArtifactComparing. RawCull injects app-specific implementations and keeps its file model, UI, sandbox policy, and culling decisions outside the package.

flowchart LR
    UI["RawCull SwiftUI and settings"] --> Integration["RawCullApplicationState composition root"]
    Integration --> AppAdapter["RawCull adapters: paths, RAW decoding, policy"]
    AppAdapter --> Contracts["PhotoAIContracts"]
    Integration --> CLIP["CoreAICLIPBackend"]
    Integration --> SAM3["CoreAISAM3Backend"]
    Integration --> Vision["VisionFeaturePrintBackend"]
    AppAdapter --> Workflows["PhotoAIWorkflows"]
    Workflows --> Contracts
    Storage["PhotoAIStorage"] --> Contracts
    CLIP --> Contracts
    SAM3 --> Contracts
    Vision --> Contracts

The arrow direction matters: PhotoAIKit does not import RawCull. A reusable package should not need to know what FileItem, RawCullViewModel, an app-specific model folder, or a burst winner means.

Responsibility Boundary

ConcernOwnerReason
Typed model, image, artifact, and segmentation contractsPhotoAIKitBackends and hosts need one stable language
Core AI CLIP and SAM 3 inferencePhotoAIKit backend productsFramework-specific tensor and inference code is reusable
Vision feature-print generation and native distancePhotoAIKit backend productThe opaque Vision payload stays behind its backend boundary
Bounded indexing, optional fallback mechanisms, segmentation, and mask selectionPhotoAIKit workflowsThese mechanisms do not depend on RawCull UI or culling policy
Optional embedding codecs and mask storesPhotoAIKit storagePersistence mechanics are reusable, but locations are not
Model installation directories and candidate orderRawCullPaths and sandbox policy belong to the host application
RAW decodingRawCullPhotoAIKit should not depend on RawParserKit or camera formats
Settings and capability wordingRawCullUser-facing state and localization belong to the app
Similarity ranking adjustments and burst groupingRawCullThese are photo-culling product decisions, not CLIP behavior
Burst-analysis cache location and lifecycleRawCullThe host owns when and where catalog results persist

Current Runtime Shape

RawCullApplicationState is the object-graph assembly boundary. It creates one RawCullAIModelRuntime, one shared SimilarityScoringModel, the focused similarity and semantic-search features, Deep Review, Qwen and Objects features, the main view model, and one RawCullIntelligenceRuntime. Identity assertions protect against accidentally constructing parallel observable state.

RawCullAIModelRuntime owns concrete providers, resource managers, Qwen inference, and separate subject/object mask stores. RawCullIntelligenceRuntime owns stable feature lifetimes and applies complete revisioned configurations from settings. See The RawCull AI Runtime for the full construction and refresh sequence. Views receive focused feature surfaces instead of the composition root or low-level scoring model:

ConsumerNarrow dependencyCapability and persistence rule
RawCullSimilarityFeatureShared SimilarityScoringModel plus RawCullSimilarityServicingOwns the public hydration, indexing, ranking, cancellation, and backend-presentation surface while persisting descriptor-valid artifacts
RawCullSemanticSearchFeatureShared scoring model and optional semantic-search serviceProjects semantic-search state, binds weakly to application selection/navigation, and never indexes missing images as a query side effect
BurstAnalysisCoordinatorSimilarity feature, scoring models, and cache repositoryOwns burst generation, progress, cache preparation, missing computation, grouping, ranking, cancellation, and derived-cache saving
DeepAIReviewControllerDeepAIReviewFeatureBuilds immutable requests from app evidence and validates the group signature before recommendations reach culling policy
RawCullAISettingsModelRawCullIntelligenceConfigurationApplyingPublishes one ordered configuration; the runtime ignores stale revisions and applies only meaningful identity changes

The safe startup and refresh path is:

  1. RawCullApp calls RawCullApplicationState.live() and retains its view model and intelligence runtime as stable @State roots.
  2. Assembly creates the shared scoring model and focused features from the initial Vision-backed configuration.
  3. RawCullAISettingsModel.refresh() asks the model runtime to validate both CLIP and both segmentation-model candidates.
  4. PhotoAIKit validates model bundles and derives model-asset fingerprints.
  5. Settings publishes a monotonically revisioned configuration. The runtime replaces similarity or semantic-search services only when their identities changed and applies segmentation selection independently.
  6. Missing, invalid, or disabled CLIP leaves burst similarity on Vision and semantic search unavailable.
  7. CLIP indexing retains valid files and logs per-file failures; Vision is not inserted into that CLIP result set.
  8. Changing the segmentation selection immediately updates the active provider. A full refresh rechecks every candidate and saved-evidence state.
stateDiagram-v2
    [*] --> VisionStartup
    VisionStartup --> CheckingModels: settings refresh
    CheckingModels --> Configuration: publish newer configuration revision
    Configuration --> CLIPSelected: preference enabled and provider ready
    Configuration --> VisionSelected: selected CLIP unavailable
    CLIPSelected --> CLIPArtifacts: keep valid per-file artifacts
    CLIPSelected --> PartialCLIP: record and exclude failed files
    VisionSelected --> VisionArtifacts: index catalog
    CheckingModels --> SegmentationSelected: activate SAM 3 or EfficientSAM

Burst similarity, semantic search, and Deep Review are separate features even when they share package code. Burst similarity may use Vision or CLIP. Semantic search requires CLIP image artifacts whose descriptor exactly matches the text provider. Deep Review uses segmentation masks and has its own availability, selection, and storage lifecycle.

Vocabulary

TermMeaning in this codebase
ProviderA backend object that performs inference or creates an artifact
Backend descriptorIdentity of the backend, model, representation, preprocessing, normalization, and configuration
Similarity artifactA descriptor plus a backend-owned payload; CLIP stores an encoded vector, while Vision stores an opaque archived observation
Source fingerprintStandardized file path, size, and modification date used to detect changed source images
Model fingerprintIdentity derived from the selected .aimodel or .aimodelc, cryptographically verified when the manifest provides a checksum
Composition rootThe one place where concrete providers, stores, paths, and app adapters are assembled
Partial CLIP resultValid CLIP artifacts plus per-file failures; failed files remain unavailable to similarity and burst grouping until a later successful index
HostThe application integrating PhotoAIKit; here, RawCull

1 - Download AI models

Download and prepare the three AI model packs

This is a terminal workbook for rebuilding RawCull’s DataComp CLIP, Meta SAM 3, and Qwen3-VL-2B-Instruct model packs from source and finishing with three Apple Managed Background Assets .aar files. It was assembled on September 26, 2026 from the local PhotoAIKit exporters, the RawCull release runbook, and the existing layout in /Users/thomas/ModelAssets/Release.

Run each numbered block in zsh and inspect the indicated output before the next block. The workbook builds in a fresh directory below /Users/thomas/ModelAssets; it does not overwrite the three existing Release/Output/*.aar files. Conversion is CPU and memory intensive and the source downloads are several gigabytes. Keep ample free space for source weights, intermediate models, three converted bundles, and three archives.

These commands produce local release candidates, not App Store approval. After packaging, compare the new hashes with the application catalog and follow the publishing runbook to update RawCull, upload the packs, and test a signed TestFlight build. A fresh conversion may produce different archive bytes even when the same source model was used.

0. What will be produced

ModelPermanent pack IDSelected installed model pathOutput archive
DataComp CLIPrawcull-clip-datacompModels/CLIP-DataCompclip-datacomp.aar
Meta SAM 3rawcull-sam3Models/SAM3sam3.aar
Qwen3-VL-2B-Instructrawcull-qwen3-vl-2bModels/Qwen/qwen3_vl_2bqwen3-vl-2b.aar

The starting evidence is the RawCull repository at /Users/thomas/GitHub/RawCull/RawCull, its sibling /Users/thomas/GitHub/RawCull/PhotoAIKit, and the notices in RawCull/ModelAssets/Notices. The old archive sizes and SHA-256 values in the application catalog describe the September 17, 2026 artifacts, not expected values for this rebuild. OpenAI CLIP and EfficientSAM are outside this three-pack release.

1. Check the Mac and select tools

Install Xcode 27 with the Core AI and Background Assets tools and select it in Xcode Settings or with xcode-select. Install git and uv if needed. The commands below merely check the selected tools:

set -e
set -o pipefail
RAWCULL_REPO='/Users/thomas/GitHub/RawCull/RawCull'
PHOTOAIKIT_REPO='/Users/thomas/GitHub/RawCull/PhotoAIKit'
MODEL_ASSETS='/Users/thomas/ModelAssets'

test -d "$RAWCULL_REPO/.git"
test -d "$PHOTOAIKIT_REPO/.git"
test -d "$MODEL_ASSETS"
command -v uv
command -v git
command -v python3
xcode-select -p
xcodebuild -version
xcrun ba-package --version

If uv is missing and Homebrew is available, install it with brew install uv, then rerun the checks. Record the exact tool versions with the eventual archive hashes. At the time this workbook was written the local Mac selected Xcode 27.0 (27A266a) and ba-package 2.0; a later toolchain can produce different output. The PhotoAIKit exporter scripts declare their Python dependencies in their own # /// script blocks, so uv run creates the appropriate isolated environments automatically. Do not combine the CLIP and SAM dependency sets by hand: their transformers requirements differ.

2. Start a clean build directory

Paste this block into the same Terminal window that will run later blocks. If you open a new window, rerun the variable definitions in this block and set BUILD_ROOT to the directory printed by echo.

set -e
set -o pipefail
RAWCULL_REPO='/Users/thomas/GitHub/RawCull/RawCull'
PHOTOAIKIT_REPO='/Users/thomas/GitHub/RawCull/PhotoAIKit'
MODEL_ASSETS='/Users/thomas/ModelAssets'
BUILD_ROOT="$(mktemp -d "$MODEL_ASSETS/Build-2026-09-26.XXXXXX")"
HF_HOME="$BUILD_ROOT/HuggingFace"
HF_HUB_CACHE="$HF_HOME/hub"
export HF_HOME HF_HUB_CACHE
mkdir -p "$BUILD_ROOT/Release/Models/Qwen" \
  "$BUILD_ROOT/Release/Notices" \
  "$BUILD_ROOT/Release/Packaging" \
  "$BUILD_ROOT/Release/Output" \
  "$BUILD_ROOT/Release/Evidence" \
  "$BUILD_ROOT/Tools"
echo "BUILD_ROOT=$BUILD_ROOT"
df -h "$MODEL_ASSETS"

This layout isolates the Hugging Face cache and the new release candidates. Never run rm -rf against the existing ModelAssets/Release tree to make room. For a later run, change the date in the mktemp template or leave it: the random suffix still creates a separate directory.

3. Freeze source and exporter revisions

The model revisions recorded in the existing RawCull catalog are:

CLIP_REV='4afec35ffe57a943d569ff7ee888061830164da8'
SAM_REV='3c879f39826c281e95690f02c7821c4de09afae7'
QWEN_REV='78448d793a7eb2f7a987a1da76d464384aa1becd'
CLIP_TOKENIZER_REV='3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268'
COREAI_MODELS_REV='475c585fdb0fe82a83c8f777f259e9414bd44c98'

COREAI_MODELS_REV is the revision pinned by the local PhotoAIKit package in the September 2026 RawCull work. These revision labels are a starting recipe. The old DataComp and Qwen provenance explicitly does not prove which exact weight bytes their exporters consumed. This rebuild records the actual source inventory so its evidence is stronger.

First save the repository states. If either repository has local changes to the exporter, review them before proceeding:

git -C "$RAWCULL_REPO" rev-parse HEAD | tee "$BUILD_ROOT/Release/Evidence/rawcull-commit.txt"
git -C "$PHOTOAIKIT_REPO" rev-parse HEAD | tee "$BUILD_ROOT/Release/Evidence/photoaikit-commit.txt"
git -C "$RAWCULL_REPO" status --short
git -C "$PHOTOAIKIT_REPO" status --short
shasum -a 256 "$PHOTOAIKIT_REPO/Tools/export_clip.py" \
  "$PHOTOAIKIT_REPO/Tools/export_sam3.py" \
  "$PHOTOAIKIT_REPO/Tools/select_sam3_asset.py" \
  | tee "$BUILD_ROOT/Release/Evidence/photoaikit-exporters-sha256.txt"

Get Apple’s converter at the selected revision in the build directory. A normal git clone also preserves its licence and Python project configuration. If the pinned revision is unavailable, stop and choose a reviewed revision; do not silently use the current main branch.

git clone https://github.com/apple/coreai-models.git "$BUILD_ROOT/Tools/coreai-models"
git -C "$BUILD_ROOT/Tools/coreai-models" checkout --detach "$COREAI_MODELS_REV"
git -C "$BUILD_ROOT/Tools/coreai-models" rev-parse HEAD \
  | tee "$BUILD_ROOT/Release/Evidence/coreai-models-commit.txt"

4. Download immutable Hugging Face snapshots

SAM 3 is gated: sign in to Hugging Face in a browser, request/receive access to facebook/sam3, accept its terms, and authenticate the terminal with uvx --from huggingface-hub hf auth login. Do not paste your access token into this workbook or a shell command. Qwen and DataComp are public, but checking their model cards and current redistribution terms is still part of release review. These downloads use Hugging Face’s CLI.

uvx --from huggingface-hub hf auth whoami
uvx --from huggingface-hub hf download \
  laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K \
  --revision "$CLIP_REV" \
  --local-dir "$BUILD_ROOT/Source/CLIP-DataComp"
uvx --from huggingface-hub hf download facebook/sam3 \
  --revision "$SAM_REV" \
  --local-dir "$BUILD_ROOT/Source/SAM3"
uvx --from huggingface-hub hf download Qwen/Qwen3-VL-2B-Instruct \
  --revision "$QWEN_REV" \
  --local-dir "$BUILD_ROOT/Source/Qwen"

The exporters do not all take a --revision or --source-dir flag. To make their ordinary model-ID lookups use the selected snapshots, populate the isolated Hugging Face cache at the same revisions and bind its main refs to those revisions. This is explicit, local to BUILD_ROOT, and leaves your ordinary Hugging Face cache alone. The CLIP exporter also obtains the OpenAI CLIP tokenizer, so cache it separately.

uvx --from huggingface-hub hf download \
  laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K \
  --revision "$CLIP_REV"
uvx --from huggingface-hub hf download facebook/sam3 --revision "$SAM_REV"
uvx --from huggingface-hub hf download Qwen/Qwen3-VL-2B-Instruct \
  --revision "$QWEN_REV"
uvx --from huggingface-hub hf download openai/clip-vit-base-patch32 \
  --revision "$CLIP_TOKENIZER_REV"

mkdir -p "$HF_HUB_CACHE/models--laion--CLIP-ViT-B-32-256x256-DataComp-s34B-b86K/refs" \
  "$HF_HUB_CACHE/models--facebook--sam3/refs" \
  "$HF_HUB_CACHE/models--Qwen--Qwen3-VL-2B-Instruct/refs" \
  "$HF_HUB_CACHE/models--openai--clip-vit-base-patch32/refs"
printf '%s' "$CLIP_REV" > "$HF_HUB_CACHE/models--laion--CLIP-ViT-B-32-256x256-DataComp-s34B-b86K/refs/main"
printf '%s' "$SAM_REV" > "$HF_HUB_CACHE/models--facebook--sam3/refs/main"
printf '%s' "$QWEN_REV" > "$HF_HUB_CACHE/models--Qwen--Qwen3-VL-2B-Instruct/refs/main"
printf '%s' "$CLIP_TOKENIZER_REV" > "$HF_HUB_CACHE/models--openai--clip-vit-base-patch32/refs/main"

The first --local-dir downloads are the inspectable evidence copies; the second downloads populate the cache actually used by the model-ID loaders. Record source hashes, including every Qwen weight shard:

find "$BUILD_ROOT/Source" -type f ! -name '.DS_Store' -print0 \
  | xargs -0 shasum -a 256 \
  | sort > "$BUILD_ROOT/Release/Evidence/source-files-sha256.txt"
rg 'safetensors|tokenizer.json|LICENSE' \
  "$BUILD_ROOT/Release/Evidence/source-files-sha256.txt" | head -80

Check that the expected source files exist before converting:

test -f "$BUILD_ROOT/Source/SAM3/model.safetensors"
test -f "$BUILD_ROOT/Source/Qwen/config.json"
find "$BUILD_ROOT/Source/Qwen" -maxdepth 1 -name '*.safetensors' -print
find "$BUILD_ROOT/Source/CLIP-DataComp" -maxdepth 1 -type f -print

DataComp binding limitation: open_clip.create_model_and_transforms uses the datacomp_s34b_b86k preset and may resolve its checkpoint through an OpenCLIP configuration rather than the checked-out LAION directory. Run it with the isolated cache and inspect the export log and cache entries. Do not claim that the LAION file was consumed solely because it was downloaded. If the resolved source differs, capture that actual source file, revision and SHA-256, or adapt the exporter to accept an explicit local weight path before release.

5. Convert DataComp CLIP

The PhotoAIKit exporter creates one two-function .aimodel plus tokenizer and bundle metadata. Its default DataComp options use ViT-B/32 at 256 px, datacomp_s34b_b86k, float16, and static shapes. The script itself pins the Python packages, including coreai-core==1.0.0b2 and coreai-torch==0.4.1.

cd "$PHOTOAIKIT_REPO"
export HF_HUB_OFFLINE=1
uv run Tools/export_clip.py \
  --model openclip-datacomp \
  --architecture ViT-B-32-256 \
  --pretrained datacomp_s34b_b86k \
  --dtype float16 \
  --output-dir "$BUILD_ROOT/Release/Models" \
  --bundle-name CLIP-DataComp \
  2>&1 | tee "$BUILD_ROOT/Release/Evidence/clip-export.log"

The exporter checks parity between the OpenCLIP tokenizer IDs and the saved PhotoAIKit tokenizer. Stop if that check fails. HF_HUB_OFFLINE=1 intentionally prevents a network fallback to an unpinned source. If OpenCLIP needs a different checkpoint repository, find its actual configured repository, pin and download it, then rerun this block in a new build directory. Do not use --overwrite until you have saved and reviewed the first result.

test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/metadata.json"
test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/tokenizer/tokenizer.json"
test -f "$BUILD_ROOT/Release/Models/CLIP-DataComp/ViT-B-32-256-datacomp_s34b_b86k_float16_static.aimodel/main.mlirb"
cat "$BUILD_ROOT/Release/Models/CLIP-DataComp/metadata.json"

6. Convert Meta SAM 3

The SAM exporter downloads the gated checkpoint by model ID, exports a source asset and an optimized runtime asset, and writes tokenizer and metadata. Select the optimized sam3_float16.aimodel so the package contains the runtime asset only. The selector refreshes the bundle fingerprint metadata.

cd "$PHOTOAIKIT_REPO"
export HF_HUB_OFFLINE=1
uv run Tools/export_sam3.py \
  --model facebook/sam3 \
  --dtype float16 \
  --output-dir "$BUILD_ROOT/Release/Models" \
  --bundle-name SAM3 \
  2>&1 | tee "$BUILD_ROOT/Release/Evidence/sam3-export.log"

python3 Tools/select_sam3_asset.py sam3_float16.aimodel \
  --bundle-dir "$BUILD_ROOT/Release/Models/SAM3"

Inspect the result. The package selector below includes only the optimized asset, tokenizer and metadata; it does not include any remaining *_source.aimodel directory.

test -f "$BUILD_ROOT/Release/Models/SAM3/metadata.json"
test -f "$BUILD_ROOT/Release/Models/SAM3/tokenizer/tokenizer.json"
test -f "$BUILD_ROOT/Release/Models/SAM3/sam3_float16.aimodel/main.mlirb"
cat "$BUILD_ROOT/Release/Models/SAM3/metadata.json"

7. Convert Qwen3-VL-2B-Instruct

Apple’s VLM exporter creates the text decoder, embedding lookup, vision encoder, tokenizer, and metadata.json in a qwen3_vl_2b directory. Use the full model: do not set --num-layers or --skip-vision. The existing RawCull bundle records a 4096 token context. The exporter has no --revision option, so the isolated cache and offline mode from section 4 matter here. Its output directory is the parent of qwen3_vl_2b.

cd "$BUILD_ROOT/Tools/coreai-models"
export HF_HUB_OFFLINE=1
uv run coreai.vlm.export --list-models
uv run coreai.vlm.export qwen3-vl \
  --max-context-length 4096 \
  --compression none \
  --output-dir "$BUILD_ROOT/Release/Models/Qwen" \
  2>&1 | tee "$BUILD_ROOT/Release/Evidence/qwen-export.log"

If the pinned converter revision does not offer qwen3-vl or one of these options, stop and inspect that revision’s --help; record and review any converter revision change. Qwen export can take a long time and use substantial memory. A process killed by macOS needs a fresh build directory or a carefully inspected incomplete-output cleanup before retrying.

QWEN_BUNDLE="$BUILD_ROOT/Release/Models/Qwen/qwen3_vl_2b"
test -f "$QWEN_BUNDLE/metadata.json"
test -f "$QWEN_BUNDLE/tokenizer/tokenizer.json"
test -f "$QWEN_BUNDLE/qwen3_vl_2b.aimodel/main.mlirb"
test -f "$QWEN_BUNDLE/embed.aimodel/main.mlirb"
test -f "$QWEN_BUNDLE/vision.aimodel/main.mlirb"
python3 -m json.tool "$QWEN_BUNDLE/metadata.json"

Check that the metadata still names Qwen/Qwen3-VL-2B-Instruct, the three assets, 448 px vision input, and 4096 context. This is a format check, not an inference test. RawCull’s Qwen provider must also load and run the bundle.

8. Stage licences, notices and build provenance

Copy the reviewed notice catalog from RawCull. These files include the complete model, tokenizer, and Apple conversion-recipe notices used by the existing release. Review their terms and dates against the newly downloaded sources before redistribution. SAM 3’s licence acceptance is required in RawCull.

for name in CLIP-DataComp SAM3 Qwen; do
  ditto "$RAWCULL_REPO/ModelAssets/Notices/$name" \
    "$BUILD_ROOT/Release/Notices/$name"
done
find "$BUILD_ROOT/Release/Notices" -type f -maxdepth 2 -print | sort

Those checked-in NOTICE.md files describe earlier published versions. Mark the staging copies as new candidates without changing their licence text:

python3 - "$BUILD_ROOT" <<'PY'
from pathlib import Path
import sys
base = Path(sys.argv[1]) / 'Release/Notices'
replacements = {
    'CLIP-DataComp': (
        'RawCull publishes this pack in the v2 model release. The release catalog\n'
        'records its download size and version, while the release host records the\n'
        'archive checksum. This notice catalog records the upstream reference revision,\n'
        'runtime fingerprint, and complete accompanying licence notices.',
        'This converted bundle is a new local release candidate. Its final archive\n'
        'size, SHA-256, Apple-assigned version, and review status must be recorded\n'
        'after packaging and upload.'),
    'SAM3': (
        'The Apple-hosted asset pack is enabled for download at the project owner\'s\n'
        'direction. Its archive byte size and SHA-256 are recorded in the external\n'
        'release evidence after packaging, while the host-correct in-pack release record\n'
        'is in `PROVENANCE.json`. This release decision does not claim an independent\n'
        'legal review. Verified licence acceptance remains required.',
        'This converted bundle is a new local release candidate. Record its archive\n'
        'size, SHA-256, Apple-assigned version, and review status after packaging\n'
        'and upload. Verified SAM licence acceptance remains required.'),
    'Qwen': (
        'RawCull publishes this pack in the v3 model release. The release catalog\n'
        'records its download size and version, while the release host records the\n'
        'archive checksum.',
        'This converted bundle is a new local release candidate. Record its final\n'
        'archive size, SHA-256, Apple-assigned version, and review status after\n'
        'packaging and upload.'),
}
for name, (old, new) in replacements.items():
    path = base / name / 'NOTICE.md'
    content = path.read_text()
    if content.count(old) != 1:
        raise RuntimeError(f'Expected release paragraph not found: {path}')
    path.write_text(content.replace(old, new))
PY

The copied PROVENANCE.json files describe the old release. Replace only the staging copies with a clearly identified record of this build. The script below preserves the licence inventory, writes the actual converted-component hashes, and points to the source inventory. It intentionally has no final .aar hash because that hash cannot be embedded inside its own archive.

python3 - "$BUILD_ROOT" "$CLIP_REV" "$SAM_REV" "$QWEN_REV" <<'PY'
import datetime, hashlib, json, pathlib, sys
root = pathlib.Path(sys.argv[1])
revisions = dict(zip(('CLIP-DataComp', 'SAM3', 'Qwen'), sys.argv[2:]))
models = {
    'CLIP-DataComp': root / 'Release/Models/CLIP-DataComp',
    'SAM3': root / 'Release/Models/SAM3',
    'Qwen': root / 'Release/Models/Qwen/qwen3_vl_2b',
}
for name, model_dir in models.items():
    notice_dir = root / 'Release/Notices' / name
    old = json.loads((notice_dir / 'PROVENANCE.json').read_text())
    components = {}
    for path in sorted(model_dir.rglob('main.mlirb')):
        digest = hashlib.sha256()
        with path.open('rb') as stream:
            for chunk in iter(lambda: stream.read(4 * 1024 * 1024), b''):
                digest.update(chunk)
        components[str(path.relative_to(model_dir))] = digest.hexdigest()
    record = {
        'catalog_version': 2,
        'release_status': 'candidate',
        'release': {
            'hosting': 'apple',
            'app_bundle_id': 'no.blogspot.RawCull',
            'asset_pack_id': {
                'CLIP-DataComp': 'rawcull-clip-datacomp',
                'SAM3': 'rawcull-sam3',
                'Qwen': 'rawcull-qwen3-vl-2b',
            }[name],
            'packaging_date': datetime.date.today().isoformat(),
            'processing_status': 'not-uploaded',
            'review_state': 'not-submitted',
        },
        'model': {
            'bundle': old.get('model', {}).get('bundle', name),
            'converted_main_mlirb_sha256': components,
        },
        'upstream': {
            'project': old.get('upstream', {}).get('project'),
            'selected_revision': revisions[name],
            'source_inventory': 'Release/Evidence/source-files-sha256.txt',
            'exporter_binding_note': 'Verify exporter log and isolated cache before claiming exact source binding.',
        },
        'conversion': {
            'photoaikit_commit_file': 'Release/Evidence/photoaikit-commit.txt',
            'coreai_models_commit_file': 'Release/Evidence/coreai-models-commit.txt',
        },
        'licences': old.get('licences', []),
    }
    (notice_dir / 'PROVENANCE.json').write_text(json.dumps(record, indent=2) + '\n')
PY

The staging provenance format is release-candidate evidence; it is not a drop-in replacement for RawCull’s checked-in PROVENANCE.json. After packaging, update the repository record using its full validated schema and the final archive hash, size, App Store Connect pack version, and processing state.

9. Create the three packaging manifests

These selector paths are relative to the current directory used by ba-package, which must be BUILD_ROOT/Release. Keep the permanent IDs and installed paths identical to RawCull’s catalog.

cat > "$BUILD_ROOT/Release/Packaging/clip-datacomp.json" <<'JSON'
{
  "assetPackID": "rawcull-clip-datacomp",
  "downloadPolicy": { "onDemand": {} },
  "fileSelectors": [
    { "file": "Models/CLIP-DataComp/metadata.json" },
    { "directory": "Models/CLIP-DataComp/tokenizer" },
    { "directory": "Models/CLIP-DataComp/ViT-B-32-256-datacomp_s34b_b86k_float16_static.aimodel" },
    { "directory": "Notices/CLIP-DataComp" }
  ],
  "platforms": ["macOS"]
}
JSON

cat > "$BUILD_ROOT/Release/Packaging/sam3.json" <<'JSON'
{
  "assetPackID": "rawcull-sam3",
  "downloadPolicy": { "onDemand": {} },
  "fileSelectors": [
    { "file": "Models/SAM3/metadata.json" },
    { "directory": "Models/SAM3/tokenizer" },
    { "directory": "Models/SAM3/sam3_float16.aimodel" },
    { "directory": "Notices/SAM3" }
  ],
  "platforms": ["macOS"]
}
JSON

cat > "$BUILD_ROOT/Release/Packaging/qwen3-vl-2b.json" <<'JSON'
{
  "assetPackID": "rawcull-qwen3-vl-2b",
  "downloadPolicy": { "onDemand": {} },
  "fileSelectors": [
    { "file": "Models/Qwen/qwen3_vl_2b/metadata.json" },
    { "directory": "Models/Qwen/qwen3_vl_2b/tokenizer" },
    { "directory": "Models/Qwen/qwen3_vl_2b/embed.aimodel" },
    { "directory": "Models/Qwen/qwen3_vl_2b/qwen3_vl_2b.aimodel" },
    { "directory": "Models/Qwen/qwen3_vl_2b/vision.aimodel" },
    { "directory": "Notices/Qwen" }
  ],
  "platforms": ["macOS"]
}
JSON

for slug in clip-datacomp sam3 qwen3-vl-2b; do
  python3 -m json.tool "$BUILD_ROOT/Release/Packaging/$slug.json" >/dev/null
done

10. Inspect and freeze the selected input files

.DS_Store files are present in the older ModelAssets/Release directories; the new candidate should not include them. Do not delete anything from the old release. Check the fresh selected directories and resolve any unexpected symlink or secret before packaging.

cd "$BUILD_ROOT/Release"
find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
  Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
  \( -name '.DS_Store' -o -type l \) -print

find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
  Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
  -type f -print0 | xargs -0 shasum -a 256 | sort \
  > Evidence/selected-inputs-sha256.txt

for slug in clip-datacomp sam3 qwen3-vl-2b; do
  xcrun ba-package evaluate "Packaging/$slug.json" \
    | tee "Evidence/$slug-evaluate.txt"
done

Read all three Evidence/*-evaluate.txt files. Each list should contain only its model bundle, tokenizer, metadata, and matching notice directory. The Qwen pack needs all three .aimodel directories. No source weight files, download cache, old archive, or other model should be selected. If the evaluation output is wrong, fix the manifest and rerun evaluation and the input inventory before packaging.

11. Build the three .aar files

ba-package creates Background Assets archives; .aar is not a ZIP file. Run it from BUILD_ROOT/Release so the relative file selectors resolve.

cd "$BUILD_ROOT/Release"
xcrun ba-package package Packaging/clip-datacomp.json \
  --output-path Output/clip-datacomp.aar --verbose \
  2>&1 | tee Evidence/clip-datacomp-package.log

xcrun ba-package package Packaging/sam3.json \
  --output-path Output/sam3.aar --verbose \
  2>&1 | tee Evidence/sam3-package.log

xcrun ba-package package Packaging/qwen3-vl-2b.json \
  --output-path Output/qwen3-vl-2b.aar --verbose \
  2>&1 | tee Evidence/qwen3-vl-2b-package.log

Do not edit an archive after this point. Any changed model, metadata, notice, or manifest requires another ba-package package run and a new hash. For the three existing App Store Connect pack records, a changed archive will become a new pack version under the same permanent ID.

12. Verify and record the result

cd "$BUILD_ROOT/Release"
for slug in clip-datacomp sam3 qwen3-vl-2b; do
  test -s "Output/$slug.aar"
  stat -f '%N|%z bytes' "Output/$slug.aar"
  shasum -a 256 "Output/$slug.aar"
  shasum -a 256 "Packaging/$slug.json"
done | tee Evidence/archive-and-manifest-sha256.txt

find Output -maxdepth 1 -type f -name '*.aar' -print | sort

The last command must print exactly these three paths:

Output/clip-datacomp.aar
Output/qwen3-vl-2b.aar
Output/sam3.aar

Compare the final selected-inputs-sha256.txt with a fresh hash pass to catch any source mutation while the archives were built:

find Models/CLIP-DataComp Models/SAM3 Models/Qwen/qwen3_vl_2b \
  Notices/CLIP-DataComp Notices/SAM3 Notices/Qwen \
  -type f -print0 | xargs -0 shasum -a 256 | sort \
  > Evidence/selected-inputs-after-sha256.txt
diff -u Evidence/selected-inputs-sha256.txt \
  Evidence/selected-inputs-after-sha256.txt

An empty diff and exit status 0 confirm that the selected inputs remained unchanged during packaging. Record the BUILD_ROOT path, Xcode version, exporter commits, source hashes, and all three archive hashes with the release candidate. The files are at:

<BUILD_ROOT>/Release/Output/clip-datacomp.aar
<BUILD_ROOT>/Release/Output/sam3.aar
<BUILD_ROOT>/Release/Output/qwen3-vl-2b.aar

13. Before uploading or calling these release files

  1. Verify that the DataComp exporter actually consumed the pinned checkpoint; the downloaded LAION snapshot alone does not prove it. For all three packs, retain actual source-weight hashes and exporter logs.
  2. Review the current model licences and all copied notice files. Keep the correct notice directory inside each .aar.
  3. Run RawCull’s make verify-model-provenance, catalog and release-metadata tests, and release preflight after updating its manifest template, Swift catalog, and checked-in provenance to the new archive values. Never leave the old archive SHA-256 in the app catalog for new .aar files.
  4. If uploading, use the existing rawcull-clip-datacomp, rawcull-sam3, and rawcull-qwen3-vl-2b App Store Connect records. Follow RawCull/Docs/newmodels.md for xcrun altool, API-key handling, processing checks, and TestFlight verification. The AppStore build must use Apple hosting and the matching App Group. A successful local .aar build does not test runtime inference.

If a step fails

SymptomCheck
401/403 downloading SAM 3Hugging Face approval and hf auth whoami; accept the gated model terms.
Offline source missingCheck the isolated HF_HUB_CACHE and the model’s refs/main; download the exact revision before re-exporting.
CLIP tokenizer parity failureDo not package; inspect the tokenizer source and OpenCLIP configuration.
SAM export leaves source assetPackage only the optimized asset selected by select_sam3_asset.py.
Qwen bundle lacks vision.aimodelRebuild with a converter that supports Qwen VLM and without --skip-vision.
ba-package evaluate lists extra filesCorrect its fileSelectors; evaluate again before packaging.
Archive hash differs from the old releaseExpected for a new conversion; update catalog and provenance before upload.

The source commands and pack layout come from PhotoAIKit’s export tools, Apple’s Core AI model recipes, Apple’s managed pack documentation, and the RawCull release runbook linked above. Verify the exact checkout used for a release: repository main branches and tool versions can change.

2 - AI Model Licence and Provenance Clearance

Current catalog status and historical model-clearance evidence.

AI model licence and provenance clearance procedure

Current code status (September 26, 2026): the 3.2.6 production catalog enables Apple-hosted DataComp CLIP, Meta SAM 3, and Qwen3-VL-2B-Instruct. The Direct/Developer ID configuration retains the historical self-hosted v3 manifest. OpenAI CLIP is excluded and EfficientSAM is not a production pack. The detailed clearance evidence below is a September 15 historical snapshot for the earlier two-pack self-hosted release; it has not been re-reviewed as a legal or provenance opinion on the current Apple-hosted Qwen pack. For current pack IDs, hashes, licences, and test-release steps, see Publishing and Testing RawCull AI Models and the application repository ModelAssets/README.md.

Technical repository evidence in the historical snapshot: 2026-09-15

Evidence record owner: Thomas Evensen, RawCull maintainer

September 26 code snapshot

Production packRecorded licenceExplicit in-app acceptanceCurrent evidence location
DataComp CLIPOpenCLIP/DataComp MIT noticeNoCatalog descriptor and ModelAssets/Notices/CLIP-DataComp
Meta SAM 3SAM License, November 19, 2025Yes, with a verified bundled textCatalog descriptor and ModelAssets/Notices/SAM3
Qwen3-VL-2B-InstructApache License 2.0NoCatalog descriptor and ModelAssets/Notices/Qwen

The three descriptors are .ready in the current code and their archive checksums are recorded in the model release guide. This table reports repository metadata, not independent legal clearance or an App Review decision. Reassess the notices and provenance for the exact packs submitted.

Historical distribution-status snapshot (September 15, 2026)

This section is a dated status snapshot. It describes the September 15 product and repository records; it is not a legal conclusion and must not be copied into a later release without a fresh evidence review.

PackCurrent product/release recordEvidence recordResidual point for the next publicationOwner/action before next publication
DataComp CLIP.ready, enabled, published in v3Archive size/SHA-256, runtime fingerprint, reference revision, tokenizer and notice hashes are recordedProvenance still records the upstream revision as a reference and leaves source_weight_sha256 nullBind the exact weight file on a rebuild or preserve a signed residual-provenance decision
OpenAI CLIP.ready in the prepared catalog, excluded from productionHistorical v2 archive and pinned source evidence remain recordedThe weight-specific licence basis described below remains a future-publication questionReassess and record a named approval before enabling it again
Meta SAM 3.ready, enabled, published in v3; verified licence acceptance requiredArchive size/SHA-256, source revision/checksum, runtime hash, complete licence and notice hashes are recordedThe upstream checkpoint is gated; the repository records the project owner’s release decision, not an independent legal opinionPreserve the decision and evidence; reopen review if terms, delivery, model, or licence text changes
EfficientSAM.blocked in the prepared catalog, excluded from productionSource/checkpoint/conversion/licence metadata are preparedFinal converted fingerprint and archive size/SHA-256 are absentKeep excluded until its descriptor and provenance pass the complete gate

The application catalog and ModelAssets records are the authoritative account of what that self-hosted v3 release contained: DataComp and SAM 3. OpenAI CLIP and EfficientSAM do not pass the inclusion flags into the production catalog or manifest template. A .ready value proves only that the product gate was opened. Model availability, a public archive, or a model-page licence badge must never be treated alone as permission for the specific conversion and redistribution.

Purpose of the reusable procedure

This document defines how RawCull clears the DataComp CLIP, OpenAI CLIP, and Meta SAM 3 model packs for public download. It covers technical provenance, licence evidence, upstream contacts, questions to ask, acceptable answers, and the final release gate.

For a new publication, do not upload a new archive, publish a new download manifest, or change a production descriptor to ready until every candidate is either:

  1. cleared under this procedure; or
  2. deliberately excluded from the release and manifest.

Existing release archives are immutable historical evidence. The safest remedy for a recorded provenance gap is a new candidate and new release tag created from pinned, hashed source files after the applicable licence decision has been reviewed; do not rewrite history by silently replacing the old record.

This is an engineering and evidence-preservation procedure, not legal advice. For an unresolved interpretation, obtain advice from a qualified lawyer who works with software copyright, open-source licensing, AI model weights, and commercial distribution in Norway and the EEA.

Reusable clearance standard

What “resolved” means

There are two independent gates.

Provenance gate

RawCull must be able to demonstrate this complete chain:

upstream owner and repository
        ↓
immutable revision and exact source-weight filename
        ↓
SHA-256 of the downloaded source file
        ↓
pinned conversion code, command, dependencies, and toolchain
        ↓
SHA-256 of the converted runtime model
        ↓
SHA-256 and byte size of the packaged Managed Background Assets pack

A filename, a local cache timestamp, or a likely upstream snapshot is not a cryptographic binding. Re-exporting from a deliberately selected and hashed source is preferable to trying to infer the origin of an old conversion.

Licence and distribution gate

RawCull must have a defensible basis for all of the following:

  • the licence applies to the exact trained weight file, not only source code;
  • conversion into an Apple Core AI model is allowed;
  • the converted derivative can be redistributed to third parties;
  • the intended distribution can be public and commercial, if RawCull is commercially distributed;
  • all notices, agreement copies, acceptance steps, use restrictions, and attribution requirements have been implemented; and
  • a gated source checkpoint may be redistributed through RawCull’s proposed delivery mechanism, if applicable.

Silence, an unanswered ticket, a community member’s assumption, widespread third-party mirroring, or the technical ability to download a file does not resolve this gate. Prefer a written response from the model owner or an authorized representative. If that is unavailable, obtain a written opinion from qualified counsel or omit the model.

Current recorded evidence

The canonical records are:

  • ModelAssets/Notices/CLIP-DataComp/PROVENANCE.json
  • ModelAssets/Notices/CLIP-OpenAI/PROVENANCE.json
  • ModelAssets/Notices/SAM3/PROVENANCE.json
  • the licence and notice files beside each provenance record

The current relevant identifiers are:

PackUpstream revision presently recordedSource-weight evidenceConverted runtime evidenceOpen issue
DataComp CLIP4afec35ffe57a943d569ff7ee888061830164da8 is a reference revision, not exporter-recorded proofExact selected source-weight SHA-256 remains null in provenanceruntime main.mlirb SHA-256 41596f6f7a9f8f8d1171b0056f4e3a90902ef88d73303713ab3bed4847b6266d; directory fingerprint 6a3639a2049b8a4ea23fe04c3083e199a4f505433f7c8bd0748b3c8d4fcb1572Bind the exact weight input on the next rebuild and recheck licence/model-card terms
OpenAI CLIP3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268, also recorded by the exporterpytorch_model.bin, SHA-256 a63082132ba4f97a80bea76823f544493bffa8082296d62d71581a4feff1576fruntime main.mlirb SHA-256 828e6ef52700c48b9c72696d785f5b45bd01a06748c538eb284cf3a42f2530da; directory fingerprint 24a20d7c5c88da2afe3ed81dca0ddf223450dd6afd1f3aff34be7acfc48f4914Preserve the named approval basis for weight redistribution; re-open the gate if that evidence is missing or changes
SAM 33c879f39826c281e95690f02c7821c4de09afae7; not exporter-boundmodel.safetensors, SHA-256 6d06f0a5f84e435071fe6603e61d0b4cc7b40e0d39d487cfd4d67d8cc11cc14aruntime main.mlirb SHA-256 43a9b88e40d193f5a6608a7fee536a78f4ba4ec5d95f1eb24db03031630f0a31Confirm whether an ungated public derivative download is compatible with Meta’s gated access flow, then rebuild with the current exporter

These hashes identify current evidence; they do not themselves approve a release.

Common provenance remediation

Perform these steps separately for every model that passes its licence gate.

  1. Choose one exact upstream repository, immutable commit, and source-weight file. Never use main, latest, an unpinned model alias, or an automatically changing download URL as the release input.
  2. Download into a new, dated evidence directory. Preserve the upstream URL, immutable revision, filename, byte size, and SHA-256 before conversion.
  3. Save the model card, licence or agreement, repository metadata, and any access terms as they appeared on the download date. Record their URLs and retrieval dates.
  4. Record the converter repository and commit, Apple coreai-models commit, coreai-core version, Python environment or package lock, conversion command, macOS version, Xcode version, and conversion timestamp.
  5. Run the conversion from that evidence directory. Do not allow the exporter to resolve or download a floating model identifier internally.
  6. Hash the complete converted model directory with the established directory-tree-sha256-v1 method and hash its runtime main.mlirb file.
  7. Validate the converted model with the same PhotoAIKit checks used by RawCull.
  8. Update the corresponding PROVENANCE.json with the actual source revision, source filename, source SHA-256, conversion command or record, tool versions, output fingerprints, and supporting evidence references.
  9. Rebuild the extensionless asset pack using the explicit selectors. Verify that the chosen runtime model, tokenizer, metadata.json, and complete notice catalog are present, and that _source.aimodel and conversion intermediates are absent.
  10. Record the new asset-pack byte size and SHA-256. The previous unpublished archive hash must not be reused for the rebuilt archive.

An example pinned Hugging Face acquisition has this form; select the correct source filename before running it:

hf download OWNER/MODEL SOURCE_WEIGHT_FILE \
  --revision IMMUTABLE_COMMIT \
  --local-dir /path/to/private/release-evidence/MODEL/source

shasum -a 256 \
  /path/to/private/release-evidence/MODEL/source/SOURCE_WEIGHT_FILE

Keep raw correspondence, access tokens, account data, and legal advice out of the public repository. Store them in private, access-controlled records. A public provenance summary may record the response date, organization, scope, and internal evidence-record identifier without publishing personal data or privileged legal advice.

DataComp CLIP clearance

Current position

RawCull currently records this pack as ready and publishes it in v3. The pinned DataComp repository page reviewed on 2026-08-22 identifies the checkpoint as MIT licensed, which is positive evidence. The remaining technical gap is that the packaged provenance does not identify and hash the exact source-weight file used by the exporter. That gap must remain visible in the decision record; the published state does not make it disappear.

Official references:

Who to contact

  1. Email LAION at contact@laion.ai. LAION publishes this address on its official legal contact page.
  2. Open a discussion on the exact Hugging Face model repository so the question and any maintainer response are tied to that checkpoint.
  3. If necessary, open an issue with the OpenCLIP maintainers to identify which upstream file the datacomp_s34b_b86k configuration resolves. OpenCLIP can help with technical identity; the model owner or counsel should resolve licence scope.

Recorded LAION email request

On 2026-08-02 at 17:20 CEST, the RawCull maintainer emailed contact@laion.ai with the subject “Written clarification requested for DataComp CLIP weight licensing and redistribution.” The request identified:

  • repository: laion/CLIP-ViT-B-32-256x256-DataComp-s34B-b86K;
  • immutable revision: 4afec35ffe57a943d569ff7ee888061830164da8;
  • proposed source-weight file: open_clip_model.safetensors;
  • model configuration: OpenCLIP ViT-B-32-256 with datacomp_s34b_b86k weights;
  • transformation: a float16 Apple Core AI runtime representation for local photo similarity and text-to-image semantic search; and
  • delivery: an optional model archive downloaded by RawCull from a public GitHub Release, including possible use with a commercially distributed version of RawCull.

The email stated that RawCull has not publicly distributed the model or its converted derivative and will not redistribute the DataComp training dataset. It proposed preserving the source byte size and SHA-256 and including the applicable MIT licence, copyright and attribution notices, model information, provenance, and checksums with the archive.

LAION was asked to confirm whether the displayed MIT licence covers the exact weight file, conversion into the proposed runtime representation, public and commercial redistribution of the derivative, and whether any additional DataComp, OpenCLIP, attribution, acceptable-use, or other conditions apply. The request also asked whether the model card’s “out of scope” deployment language is safety guidance or an additional legal restriction.

Status: awaiting a substantive response from LAION or another person authorized to clarify the rights applicable to the weights. Sending the email does not clear the pack. Preserve the original message and any response in the private evidence register; do not add the maintainer’s personal email address or complete mail headers to this public repository.

Questions to ask

Identify the exact repository, revision, and proposed source file, then ask:

  1. Is that exact trained weight file offered under the MIT licence shown on the model repository?
  2. Does that permission cover conversion into another runtime representation and redistribution of the converted weights with a desktop application?
  3. Is public and commercial redistribution permitted, provided the MIT notice is included?
  4. Are any DataComp dataset terms, OpenCLIP terms, attribution requirements, or use restrictions additional to the displayed MIT licence applicable to the weight file?

Required technical work

  1. Select exactly one of the upstream weight files rather than leaving the exporter to resolve a model alias.
  2. Download it at revision 4afec35ffe57a943d569ff7ee888061830164da8 and record its SHA-256 and byte size.
  3. Re-export DataComp CLIP from that local file under the common procedure.
  4. Replace the null source revision and source checksum in ModelAssets/Notices/CLIP-DataComp/PROVENANCE.json.

Sufficient resolution

The pack may pass this gate when both conditions hold:

  • the new conversion has a complete pinned-and-hashed provenance chain; and
  • LAION or another demonstrably authorized model owner confirms the licence scope in writing, or qualified counsel concludes in writing that the repository’s MIT designation and accompanying materials are sufficient for the intended distribution.

If the answer is negative or remains materially ambiguous, replace the model with a checkpoint having explicit weight-level redistribution terms or omit the DataComp pack.

DataComp CLIP - no answer

If LAION does not provide a substantive answer after a documented follow-up, silence neither grants additional permission nor withdraws the permission already stated in the published materials. A release under the displayed MIT licence would rely on the public licence evidence rather than individualized clearance from LAION. Record it as a maintainer risk-acceptance decision, not as an upstream-approved or upstream-cleared release.

The evidence supporting that decision is:

  • the exact pinned model repository identifies the model as License: mit and contains the weight files;
  • Hugging Face’s licence documentation describes model-card licence metadata as communicating the permissions attributed to repository content; and
  • the MIT licence permits use, modification, publication, redistribution, sublicensing, and sale when its copyright and permission notice accompanies copies or substantial portions.

The residual ambiguity is that the model repository has MIT metadata but no standalone LICENSE file identifying the trained weights and their copyright holder. The OpenCLIP MIT notice clearly covers the OpenCLIP software, but an unanswered inquiry leaves no individualized confirmation that it is also the intended notice for the trained weights and converted derivative. The model card’s “out of scope” deployment language appears as safety and intended-use guidance rather than licence text, but this interpretation has not been confirmed by LAION. Qualified Norwegian counsel remains the recommended way to resolve that ambiguity before a public or commercial release.

If the maintainer nevertheless decides to release without an answer, complete all of the following before publication:

  1. Send and preserve one documented follow-up to LAION. Record the original request, follow-up date, response deadline, and absence of a substantive answer in the private evidence register.
  2. Select open_clip_model.safetensors from revision 4afec35ffe57a943d569ff7ee888061830164da8 as the only conversion input. Its pinned Hugging Face metadata reports a byte size of 605189364 and SHA-256 92c26d60d3200ed5ed040dff31a8d19f8140648da8007216c25744c478deef27. Download the file independently and verify both values before conversion.
  3. Re-export the Apple Core AI model from that pinned local file. Do not merely add the upstream hash to the provenance record for the existing conversion; the released derivative must be cryptographically tied to the verified source file.
  4. Preserve dated copies of the pinned repository tree, model card, repository API metadata, MIT licence, OpenCLIP notice, conversion command, dependency versions, and all input and output checksums.
  5. Package the complete applicable MIT copyright and permission notice, OpenCLIP and tokenizer notices, model card, safety limitations, provenance, and conversion information with every redistributed model archive. A link alone is not a substitute for including the required notice.
  6. Keep DataComp CLIP identified as a separate third-party model asset. RawCull’s own MIT licence does not relicense the model, and RawCull must not claim ownership of or permission to redistribute the DataComp training dataset.
  7. Complete the common provenance procedure, PhotoAIKit validation, asset-pack inspection, archive hashing, manifest verification, and download tests. Change PROVENANCE.json and the production catalogue to ready only after those technical controls describe the new release candidate accurately.
  8. Add a signed and dated release decision stating the evidence relied upon, the unresolved licence ambiguity, whether RawCull is free or commercial, the intended distribution countries and channels, and the responsible maintainer’s acceptance of the residual risk.
  9. Keep the model pack independently removable from the manifest and release so distribution can be suspended promptly if a credible rights claim or contrary clarification is received.

Releasing after these steps may provide a defensible MIT-compliance position, but it does not eliminate the residual legal risk created by the absence of a weight-specific licence file or an authorized response. This procedure records the evidence and decision; it is not a legal opinion.

OpenAI CLIP clearance

Current position

RawCull retains this pack as ready in the prepared catalog, but excludes it from the production catalog and v3 manifest. The historical v2 record remains relevant if this model is considered for a future release. The OpenAI CLIP source repository contains an MIT licence covering the software and associated documentation. The Hugging Face checkpoint page reviewed on 2026-08-22 still does not display a clear weight-level licence designation. A community assumption that the repository licence covers the weights is not authoritative evidence.

The model card also characterizes deployed uses as out of scope. Clarify whether this is safety guidance or an enforceable distribution/use condition, and independently assess RawCull’s use, testing, limitations, and disclosures.

Official references:

Who to contact

  1. Contact OpenAI Support using the chat control at help.openai.com. Ask that the request be routed to the team responsible for CLIP/open-source model licensing. Retain the ticket or conversation identifier.
  2. Submit the same narrowly framed question through the feedback form linked by the official CLIP model card.
  3. Open a model-specific discussion on the Hugging Face Community page using its New discussion action. This requires a Hugging Face login. Use a discussion rather than a pull request so the licensing question remains attached to the exact model distribution.
  4. A response from an OpenAI employee or repository maintainer authorized to address the model’s licence is preferred. An unsupported answer from another community member is not sufficient.

GitHub issue creation for openai/CLIP is currently restricted, so it should not be the only planned contact route.

Recorded support outcome and Hugging Face escalation

OpenAI Support declined to confirm or provide an authoritative interpretation that the MIT licence in openai/CLIP applies to the pretrained weight files in openai/clip-vit-base-patch32. Support also declined to confirm that the licence permits RawCull’s intended commercial redistribution. The response said that the applicable terms must be determined from the licence and notices shipped with the exact code and weights being used.

This historical response did not by itself clear the pack. RawCull’s later product record nevertheless marks the pack ready, so the private decision register must identify the subsequent evidence, responsible approver, date, and accepted scope. If it cannot, treat that as a blocker for the next publication rather than pretending the historical uncertainty was resolved. The conclusion at the time of the support exchange was:

CLIP source code: MIT licensed.
openai/clip-vit-base-patch32 pretrained weights: no explicit weight-specific
licence identified in the downloaded distribution, and OpenAI Support did not
confirm that the source-repository MIT licence applies.
Commercial redistribution: unresolved and blocked pending sufficient evidence.

The OpenAI Report Content form is intended for reports of potentially illegal or policy-violating content. It is not evidence of permission and should not be treated as the route for prospective licensing clearance. Support’s recommended next step was to ask the maintainers on the distribution source. For Hugging Face, use the model-specific Community page linked above rather than the general Hugging Face forum.

On 2026-08-02, the RawCull maintainer opened Hugging Face discussion #72, “License applicable to pretrained CLIP ViT-B/32 weights,” asking the OpenAI maintainers to identify the licence covering the hosted weights and to confirm whether it permits commercial use and redistribution. The request also asks the maintainers to add the applicable licence identifier or licence file to the model repository. Opening the discussion records the escalation but does not clear the pack; retain the response and assess the responder’s authority and the scope of any answer before changing the release decision.

Suggested discussion title:

Licence applicable to pretrained CLIP ViT-B/32 weights

Suggested discussion body:

Could the OpenAI maintainers clarify the licence applicable specifically to
the pretrained weight files in openai/clip-vit-base-patch32?

In particular, does OpenAI intend the MIT License from the official
openai/CLIP repository to cover the original ViT-B/32 checkpoint and the
converted weight files hosted here, including commercial use, conversion to
another runtime representation, and redistribution subject to the MIT
conditions?

If so, could you add the applicable licence identifier and/or LICENSE file to
this model repository so downstream users can establish a reliable licensing
record?

A maintainer comment may clarify intent, but the strongest resolution is a licence file, model-card licence declaration, or written statement from the rights holder that expressly covers the exact weights and proposed use. Until a responsible approver records the basis that supersedes this historical finding, do not use existing v2 availability as the sole basis for a new commercial redistribution decision.

Questions to ask

Provide this exact identity:

  • repository: openai/clip-vit-base-patch32;
  • revision: 3d74acf9a28c67741b2f4f2ea7635f0aaf6f0268;
  • file: pytorch_model.bin;
  • SHA-256: a63082132ba4f97a80bea76823f544493bffa8082296d62d71581a4feff1576f.

Ask:

  1. What licence governs this exact trained weight file?
  2. Does the OpenAI CLIP MIT licence apply to it, or is there a separate licence or set of terms?
  3. May RawCull convert the weights into an Apple Core AI representation and publicly redistribute that converted derivative as an optional desktop-app download, including in a commercial application?
  4. Which copyright notice, attribution, model card, use limitation, or other terms must accompany the derivative?
  5. Is the model card’s statement that deployed uses are out of scope guidance, or does OpenAI intend it as a legal restriction on deployment or redistribution?

Required technical work

After licence clearance, re-export from the exact local source file and revision above. Record the conversion command and bind the new output to the source SHA-256. Update the provenance record and rebuild the extensionless asset pack.

Sufficient resolution

The pack may pass this gate only with:

  • an authoritative written statement identifying terms that permit the exact intended conversion and redistribution; or
  • a written legal opinion accepting a clearly documented alternative basis.

A response that merely restates the code repository’s MIT licence without addressing the trained weights is not sufficient. If OpenAI does not answer or the answer does not permit the proposed distribution, omit OpenAI CLIP or use a replacement checkpoint with explicit weight-level terms.

Meta SAM 3 clearance

Current position

The SAM License dated November 19, 2025 defines the trained weights as SAM Materials and grants rights to use, reproduce, distribute, modify, and create derivatives. Distribution must remain under that agreement and a copy of the agreement must accompany the materials. RawCull now packages the complete agreement and requires acceptance of its verified text.

The separate unresolved question arises because Meta’s official checkpoint is gated. The model page requires a user to log in and share contact information, and Meta’s repository instructs users to request access and authenticate before downloading. Hugging Face documents that a gated model’s authors control access. The licence text appears permissive about redistribution, but a public GitHub download would let downstream users obtain the derivative without going through Meta’s upstream access flow.

The repository’s current product decision nevertheless marks SAM 3 ready, enables it in production v3, and records verified archive metadata. This technical state does not erase the review topic below or claim independent legal advice; preserve the decision record and reopen it when the model, licence, delivery mechanism, or upstream terms change.

Official references:

Who to contact

  1. Email Meta Open Source at opensource@meta.com. Meta publishes this contact address on its official Open Source terms page. Ask that the request be routed to the SAM 3 model/licensing owner.
  2. Open a narrowly scoped issue in facebookresearch/sam3 and/or a discussion on facebook/sam3. Link the public question from the private email so the project team can answer in its preferred channel.
  3. Contact Hugging Face only for clarification of how its gate operates. Hugging Face is the hosting platform; its support cannot substitute for permission or interpretation from Meta as the model owner.
  4. If Meta does not give a clear response, obtain a written opinion from qualified counsel or omit SAM 3 from public hosting.

Questions to ask Meta

Provide this exact identity and delivery proposal:

  • repository: facebook/sam3;
  • revision: 3c879f39826c281e95690f02c7821c4de09afae7;
  • file: model.safetensors;
  • SHA-256: 6d06f0a5f84e435071fe6603e61d0b4cc7b40e0d39d487cfd4d67d8cc11cc14a;
  • transformation: float16 Apple Core AI derivative for local inference;
  • delivery: optional RawCull Managed Background Asset from a public GitHub Release;
  • licence handling: complete SAM License packaged with the derivative and explicit in-app acceptance tied to the licence-text SHA-256.

Ask:

  1. Does section 1 of the SAM License permit this converted derivative to be distributed from a public, ungated GitHub Release?
  2. Must every downstream RawCull user independently request access through Meta’s Hugging Face gate before receiving the converted derivative?
  3. If downstream gating is required, what information and approval must RawCull collect, and may RawCull technically administer that gate?
  4. Is packaging the complete agreement and requiring verified in-app acceptance sufficient to distribute under the same agreement?
  5. Are there additional attribution, branding, reporting, geographic, trade control, or prohibited-use measures RawCull must implement?
  6. Does the answer apply to both free and commercial distribution of RawCull?

Possible outcomes

Meta confirms public ungated redistribution. Preserve the response, comply with every stated condition, re-export from the pinned source, update the provenance record, and have counsel review any material ambiguity before release.

Meta requires each recipient to pass its gate. Do not publish SAM 3 as a public GitHub Release asset. Managed Background Assets with an anonymous public origin will not reproduce individualized Hugging Face approval. Omit SAM 3 or design a separate authenticated acquisition flow and review it independently.

Meta declines, gives an unclear answer, or does not respond. Keep SAM 3 blocked. Obtain a legal opinion or omit it. Do not treat the licence’s general redistribution wording as resolving a separately imposed access condition without an accountable decision.

Required technical work after clearance

Re-export from the recorded revision and model.safetensors SHA-256, recording the complete conversion command and environment. Preserve the full SAM License with the pack, maintain explicit acceptance against its checksum, and rebuild the extensionless asset pack. Re-check the official licence immediately before publication because the agreement states that Meta may modify it.

Contact-request template

Use a real name and reply-capable address. Do not send an archive, proprietary source, access token, or confidential material unless the recipient requests it through an appropriate channel.

Subject: Written clarification requested for redistribution of [MODEL] in RawCull

Hello,

I maintain RawCull, a macOS photo-culling application:
https://github.com/rsyncOSX/RawCull

I have not published or uploaded the model derivative described below. I am
seeking written clarification before doing so.

Upstream repository: [URL]
Immutable revision: [COMMIT]
Source weight file: [FILENAME]
Source SHA-256: [SHA256]

RawCull converts this file to a float16 Apple Core AI model for local inference.
The proposed delivery is an optional Managed Background Asset downloaded from a
public GitHub Release. The archive will include the applicable complete licence,
notices, attribution, provenance, and checksums. [Describe explicit acceptance,
if applicable.] RawCull is [free/paid; choose the accurate description].

Please confirm:

1. Which licence or terms govern this exact trained weight file?
2. May it be converted into this runtime representation?
3. May the converted derivative be publicly redistributed in this manner,
   including as part of a commercial application if applicable?
4. What notices, acceptance flow, attribution, gating, use restrictions, or
   other conditions must RawCull implement?

I would appreciate a response from the model owner or a person authorized to
clarify its licensing and distribution conditions. If another team handles
this request, please route it or identify the correct contact.

Thank you,
[NAME]
[CONTACT DETAILS]

For SAM 3, append the six SAM-specific questions above. For OpenAI CLIP, append the exact question about whether the MIT licence covers the trained weights and whether the model card’s deployment language is guidance or a restriction. For DataComp, ask whether the repository’s MIT designation covers the exact selected file.

When to involve a lawyer

Engage a Norwegian lawyer before publication if any of these apply:

  • the model owner does not answer;
  • the answer is informal, conditional, or ambiguous;
  • a licence permits redistribution but an upstream gate suggests a different access policy;
  • commercial use, downstream acceptance, trade controls, prohibited uses, or indemnification requires interpretation;
  • the responder’s authority to bind the rights holder is uncertain; or
  • RawCull will distribute in multiple jurisdictions.

Look for counsel with experience in copyright, software and open-source licensing, technology transactions, AI model weights, EEA law, and US/EU trade controls. The Norwegian Bar Association member search can verify professional membership. The Norwegian Industrial Property Office also publishes guidance on choosing an IP adviser.

Give counsel a private evidence bundle containing:

  • the exact model card and licence captured on the download date;
  • upstream repository, revision, source filename, SHA-256, and download method;
  • the conversion recipe and explanation of what the derivative contains;
  • complete notice catalogs;
  • every upstream response with message headers, dates, ticket IDs, and links;
  • RawCull’s intended countries, free or paid status, distribution channels, end-user flow, and licence-acceptance UI; and
  • the proposed GitHub release and Managed Background Assets architecture.

Ask for a written conclusion for each model covering conversion, redistribution, commercial use, notices, downstream terms, gating, and any required technical controls. Keep privileged advice private. Record only the resulting release decision and non-confidential obligations in the repository, unless counsel approves broader disclosure.

Evidence and decision record

Maintain a private register with at least these fields:

FieldRequired content
ModelExact upstream owner and repository
Source identityImmutable revision, filename, byte size, SHA-256
Licence identityName/version, official URL, captured file SHA-256, retrieval date
Contact recordOrganization, channel, date, ticket/issue ID, responder and stated authority
Permission scopeConversion, derivative redistribution, public access, commercial use, territories
ConditionsNotices, attribution, acceptance, gating, use restrictions, trade controls
Legal reviewCounsel, date, private matter/reference number, approved/blocked conclusion
ConversionScript/commit, command, dependencies, toolchain, timestamp, output hashes
PackExplicit selector manifest, asset-pack byte size and SHA-256, notice verification
DecisionReady, blocked, replaced, or omitted; responsible approver and date

Do not mark a contact item complete merely because a message was sent. Record the actual answer and whether it addresses the exact file and proposed delivery.

Final release checklist

A pack can change from blocked to ready only when every applicable item is complete:

  • Exact upstream owner, repository, immutable revision, and source file are recorded.
  • Source byte size and SHA-256 were computed before conversion.
  • The licence is captured from an official source and its checksum is recorded.
  • The licence demonstrably covers the trained weights and converted derivative.
  • Public and commercial redistribution is permitted for RawCull’s actual delivery model.
  • Any gated-access question is answered by the owner or resolved by qualified counsel.
  • All required notices, agreement copies, attribution, acceptance, and use controls are implemented.
  • The model was re-exported from the pinned local source under a recorded command and environment.
  • Converted output fingerprints and runtime hashes are recorded.
  • PhotoAIKit validation passes.
  • A new extensionless Managed Background Assets pack was built with explicit selectors and inspected.
  • The asset-pack byte size and SHA-256 are recorded.
  • PROVENANCE.json, NOTICE.md, and the application catalogue agree.
  • A responsible human has signed and dated the release decision.

If even one required item remains open, keep that pack blocked or omit it.

Publication sequence after clearance

RawCull may publish a ready subset. Every candidate must still have an explicit ready, blocked, replaced, or omitted decision, and every blocked/omitted asset must be absent from that manifest. The current v3 record follows this rule by publishing DataComp CLIP and SAM 3 while excluding OpenAI CLIP and EfficientSAM.

After the decisions are complete:

  1. Re-export every approved model from its pinned and hashed source.
  2. Update the notice catalogs and change only genuinely approved catalogue descriptors to ready.
  3. Rebuild and inspect the extensionless asset packs; record their new hashes and sizes.
  4. Generate and inspect the self-hosted download manifest with a non-beta or corrected ba-package toolchain.
  5. Create the dedicated RawCull-AI-Models release as a draft and upload only approved packs, their required remote asset names, the manifest, and public evidence.
  6. Verify every URL, redirect, byte size, and checksum while authenticated to the draft if necessary.
  7. Run download, acceptance, validation, removal, and licence-change tests against a non-production environment.
  8. Publish the release only after a final human review confirms that the manifest contains no blocked model and all obligations are satisfied.

Official-source summary

The following official sources were rechecked on 2026-08-22:

At that review date the DataComp page displayed License: mit; the OpenAI checkpoint page did not display a corresponding weight-specific licence label; and the official SAM 3 checkpoint remained gated with the SAM License linked. These observations are evidence inputs, not legal conclusions.

Recheck all licence text, model-card metadata, gates, contacts, and official links immediately before release because they can change.

3 - AI Models in RawCull

Code-level guide to local CLIP, Vision, SAM 3, Qwen, Deep Review, and numbered Objects analysis in RawCull.

AI Models in RawCull

RawCull uses several local machine-learning backends, but it does not treat them as interchangeable. Each model family has a deliberately narrow job:

Model or backendRawCull jobOutput used by RawCull
DataComp CLIPImage similarity, burst grouping, semantic search, and coarse subject labels for Deep ReviewNormalized image/text embedding vectors and cosine distances/similarities
OpenAI CLIPFully implemented alternative CLIP bundle; currently excluded from the production model listThe same typed CLIP artifacts as DataComp, with a different model fingerprint
SAM 3Prompted subject segmentation for Deep Review and separate instance segmentation for ObjectsA chosen subject mask, or up to eight numbered masks per concept
Qwen3-VL-2B-InstructStandalone photo assessment; concept discovery and board interpretation in ObjectsA photo assessment, validated object concepts and per-object findings, or a visible retryable response failure
Apple Vision feature printAlways-available image-similarity fallbackOpaque Vision feature-print artifacts and native distances

All inference stays in the application process. The downloadable model assets are installed separately because they are large, but the analysis path does not send photographs to a remote inference service.

This page was checked against the local RawCull version-3.2.6 working tree on September 26, 2026, including the packaged-model evaluation and later in-app Objects observations. RawCull source links follow that branch. PhotoAIKit links use revision 77cc1d84, pinned by this checkout. The local working tree may be ahead of the published branch until those changes are pushed.

Source Catalog: RawCull/Intelligence

The Intelligence directory is organized by responsibility rather than by one model per folder. Start with this catalog when tracing the code.

Composition and contracts

SourceResponsibility
Composition/RawCullAIModelRuntime.swiftOwns model-resource managers, validated CLIP and SAM 3 providers, the Vision fallback, Qwen inference runtime, mask stores, and the active segmentation pipeline.
Composition/RawCullIntelligenceRuntime.swiftAssembles the complete application AI graph and preserves stable feature identities while services change.
Contracts/RawCullAIModels.swiftDefines model choices, paths, capability states, and saved-artifact evidence.

Model management

SourceResponsibility
RawCullAIModelDownloadCatalog.swiftProduction model inventory, inclusion switches, asset-pack identifiers, versions, byte counts, checksums, licences, and provenance links.
RawCullAIModelDownloadService.swiftBackground Assets download coordination and installed-location resolution.
RawCullAIModelDownloadsModel.swiftObservable download, licence, progress, removal, and installed-location state.
RawCullAIModelResourceManager.swiftActor-isolated validation and provider construction with a metadata snapshot cache.
RawCullAISettingsModel.swiftApplies installed locations, refreshes capabilities, stores user selections, and publishes revisioned runtime configurations.
SourceResponsibility
RawCullVisionSimilarityService.swiftDefines the shared similarity-service boundary, Vision implementation, CLIP implementation, RAW decoding adapter, finite-vector recovery, and artifact validation.
SimilarityScoringModel.swiftOwns indexed artifacts, hydration, persistence, image ranking, grouping, semantic-search state, and CLIP-based subject classification.
RawCullSimilarityFeature.swiftStable application-facing similarity surface with cancellation and generation gates.
RawCullSemanticSearchService.swiftEncodes a text query, admits compatible CLIP artifacts, compares image and text vectors, and ranks deterministically.
RawCullSemanticSearchFeature.swiftPresentation and application-target adapter for semantic search.

Deep Review with SAM 3 and CLIP

SourceResponsibility
DeepAIReviewController.swiftConverts the current RawCull selection, burst evidence, sharpness evidence, subject label, and AF point into a review request.
DeepAIReviewFeature.swiftOwns review state and implements the complete decode → prompt → mask → score → recommendation pipeline.
SubjectMaskFocusScorer.swiftComputes subject-only broad, local, and fine detail evidence from an image and SAM mask.
DeepAIReviewMaskOutlineRenderer.swiftTurns a persisted filled mask into a display outline.

Objects: Qwen discovery, SAM 3 instances, Qwen review

SourceResponsibility
ObjectAnalysis/RawCullObjectAnalysisFeature.swiftBatch coordination, availability, cancellation, retry, private capture, and stage timings.
ObjectAnalysis/ObjectConceptDiscovery.swiftAutomatic prompt, concept validation, and Specific Concepts parsing.
ObjectAnalysis/ObjectInstanceDeduplicator.swiftFilters weak masks, merges near-identical masks across concepts, and assigns board IDs.
ObjectAnalysis/ObjectReviewBoardRenderer.swiftRenders the 2,048-pixel overview and numbered, outlined crops for Qwen; fails if a numbered crop cannot be prepared.
ObjectAnalysis/ObjectJSONEnvelope.swift and ObjectAnalysis/ObjectAnalysisResponseDecoder.swiftRecover one JSON object from a wrapper, then validate fields, confidence, list limits, and board IDs.
ObjectAnalysis/ObjectAnalysisModels.swiftMode, instance, assessment, progress, timing, and result types.
ObjectAnalysis/ObjectMaskOutlineRenderer.swiftDetail-view contour from a stored grayscale instance mask.
Views/AIAnalysis/ObjectAnalysisView.swiftControls, status table, numbered overlays, crop, per-object detail, and retry.

The object-set workflow uses PhotoAIKit’s ObjectSegmentationService, ObjectMaskMemoryStore, optional ObjectMaskDiskStore, and SAM 3 ObjectInstanceSegmenting contract. Its cache is separate from the Deep Review subject-mask cache.

Qwen

SourceResponsibility
QwenInferenceRuntime.swiftActor-owned Qwen provider validation, lazy vision-language model loading, session creation, prompt construction, response decoding, and invalidation.
RawCullQwenAnalysisFeature.swiftMain-actor batch operation, image loading, progress, per-file failure isolation, result retention, and cancellation.
QwenPhotoAssessment.swiftStructured response schema, validation, free-form fallback, aggregate score, and result types.

Persistence and burst consumption

PerFileAnalysisArtifactStore persists descriptor-bearing similarity artifacts. The burst-analysis files consume those artifacts, build groups, cache results, and reject incompatible cache data. They are not model runtimes themselves. This separation is important: a CLIP model produces an embedding; RawCull’s burst policy decides what that embedding means for grouping and culling.

The Model Inventory Shipped by RawCull

The authoritative inventory is RawCullAIModelDownloadCatalog.prepared. production filters that inventory through code-only inclusion switches.

Production modelAsset-pack IDInstalled model path inside packDownload sizeInstalled size
DataComp CLIP, ViT-B/32 at 256 pxrawcull-clip-datacompModels/CLIP-DataComp282,967,354 bytes307,800,172 bytes
Meta SAM 3rawcull-sam3Models/SAM31,542,689,931 bytes1,667,570,378 bytes
Qwen3-VL-2B-Instructrawcull-qwen3-vl-2bModels/Qwen/qwen3_vl_2b3,754,599,603 bytes5,395,195,663 bytes

OpenAI CLIP exists in RawCullCLIPModel, has a resource manager, and can be selected by the runtime, but includeOpenAICLIP is currently false. DataComp CLIP, SAM 3, and Qwen download are enabled. SAM 3 requires explicit acceptance of its bundled, hash-verified licence before download; the DataComp and Qwen licences do not require an extra acceptance action.

How AI Modules Are Instantiated

The application has two stable roots: RawCullViewModel for general app state and RawCullIntelligenceRuntime for AI-facing state. They are created once in RawCullApp.init() and retained in SwiftUI @State.

flowchart TD
    App["RawCullApp.init()"] --> State["RawCullApplicationState.live()"]
    State --> Models["RawCullAIModelRuntime"]
    State --> Downloads["RawCullAIModelDownloadsModel"]
    State --> QwenFeature["RawCullQwenAnalysisFeature"]
    State --> Objects["RawCullObjectAnalysisFeature"]
    State --> DeepFeature["DeepAIReviewFeature"]
    State --> Settings["RawCullAISettingsModel"]
    State --> Scoring["SimilarityScoringModel"]
    Scoring --> Similarity["RawCullSimilarityFeature"]
    Scoring --> Semantic["RawCullSemanticSearchFeature"]
    DeepFeature --> Controller["DeepAIReviewController"]
    State --> VM["RawCullViewModel"]
    State --> Runtime["RawCullIntelligenceRuntime"]
    Runtime --> Models
    Runtime --> QwenFeature
    Runtime --> Objects
    Runtime --> Similarity
    Runtime --> Semantic
    Runtime --> Controller
    Runtime --> Settings

The exact construction order in RawCullApplicationState.make is significant:

  1. RawCullAIModelRuntime is supplied by live(). Its initializer creates resource-manager actors for SAM 3 and both CLIP choices, one QwenInferenceRuntime, the Vision provider/service, mask stores, and a placeholder unavailable segmentation pipeline.
  2. RawCullAIModelDownloadsModel is created with the production catalog and application paths.
  3. RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature receive the same Qwen inference actor. Objects also receives its memory and optional disk instance-mask stores.
  4. DeepAIReviewFeature starts with the model runtime’s current segmentation capability. It is bound back to the model runtime so a later SAM provider can install a real pipeline without replacing the feature.
  5. RawCullAISettingsModel receives the model runtime, downloads model, Qwen feature, preferences store, and saved-evidence scanner.
  6. Settings produces a synchronous initial configuration. Before asynchronous validation finishes this normally selects the Vision fallback.
  7. A single SimilarityScoringModel is created. Both similarity and semantic search share this same artifact/state owner.
  8. Stable feature and controller objects are created around those models.
  9. RawCullViewModel receives the exact same feature objects.
  10. RawCullIntelligenceRuntime retains the graph and binds the narrow weak application contexts.
  11. Settings binds its weak configuration consumer and immediately publishes the first revision.

Debug assertions verify identity sharing. These checks are not cosmetic: a second Qwen inference actor, scoring model, Objects feature, or Deep Review feature would split model state, tasks, caches, and UI observation.

The first asynchronous validation begins from the main view’s .task:

.task {
    await intelligenceRuntime.settingsModel.refresh()
}

refresh() asks the downloads model for an installed-location snapshot. That snapshot flows through settings to RawCullAIModelRuntime, which validates Qwen and refreshes CLIP and SAM 3 capabilities. See The RawCull AI Runtime for the concrete PhotoAIKit provider handoff, feature wiring, lifetime, and reconfiguration path. In particular, the download snapshot supplies URLs; PhotoAIKit factories validate bundles and create typed providers; the model runtime retains those providers; and Settings sends selected services in a revisioned configuration to the stable intelligence runtime. Qwen follows its own actor path and updates its existing analysis feature through model status.

Why Qwen has its own inference runtime

It may look simpler to put Qwen’s provider, loaded model, and inference methods directly inside RawCullAIModelRuntime. The two types have different jobs, however, and keeping those jobs separate makes their concurrency and lifetimes clear.

Think of RawCullAIModelRuntime as the coordinator for the application’s model room. It knows which resources are installed, validates capabilities, selects the services that RawCull should expose, and publishes those choices on the main actor. QwenInferenceRuntime, by contrast, is the specialist operating one machine in that room. Its actor protects Qwen-specific mutable state: the validated provider, the lazily loaded vision-language model, and the generation counter used to reject work from a model that has since been removed or replaced. It also owns Qwen-specific work such as creating sessions, building prompts, running generation, and decoding responses.

This boundary matters because Qwen inference can suspend for comparatively long operations such as loading the model and generating a response. Those operations should be serialized by the Qwen actor without turning the main-actor RawCullAIModelRuntime into the place where heavy inference runs. It also keeps Qwen’s two-stage lifecycle—validate a lightweight provider now, then load the heavy model only when it is first used—independent of the CLIP and SAM 3 resource lifecycles.

Separate does not mean unrelated. RawCullAIModelRuntime.init creates one QwenInferenceRuntime and retains it as qwenInference. During application assembly, that exact instance is passed to RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature. Consequently, there is one owner of Qwen’s provider and loaded model, while the model runtime remains the composition point that creates and coordinates the application’s complete collection of AI backends. In short:

  • RawCullAIModelRuntime answers which AI capabilities are available and how they fit into the application;
  • QwenInferenceRuntime answers how one Qwen request is safely executed; and
  • constructing the latter inside the former guarantees a single, shared Qwen runtime rather than independent copies with competing model state.

Model Discovery, Validation, and Provider Construction

CLIP and SAM 3 use one RawCullAIModelResourceManager<Provider> actor per model choice. The actor owns:

  • caller-ordered fallback candidate URLs;
  • the current managed asset-pack URL;
  • the PhotoAIKit ModelProviderFactory;
  • a lightweight file-metadata snapshot; and
  • the cached capability/provider result.

Changing the managed URL clears the cache. load() prepends the managed URL to any fallback candidates, snapshots every directory entry, and returns its cached result if path, kind, size, modification date, and resolved symlink path have not changed. When the snapshot changes, PhotoAIKit remains authoritative: it checks metadata.json, the declared model asset, required tokenizer files, accepted .aimodel/.aimodelc extensions, and the model fingerprint or manifest checksum. Only then does the factory create the concrete provider.

The file-metadata snapshot is an optimization, not a security decision. It decides when full validation may be reused; it does not replace PhotoAIKit’s model-bundle validation.

Qwen has a separate actor because its lifecycle differs. Its validation builds an immutable CoreAIQwenProvider; its much heavier CoreAIVisionLanguageModel is created lazily on the first assessment and reused across later LanguageModelSession values.

How CLIP Works in RawCull

CLIP places images and text in a shared vector space. RawCull uses that property in three ways: image-to-image distance, text-to-image semantic search, and a small closed-set subject-label pass that helps choose SAM prompts.

Provider construction and identity

PhotoAIKit’s CoreAICLIPProvider is an actor implementing image embedding, artifact generation/comparison, text embedding, and image/text comparison. During initialization it:

  1. validates the supplied model bundle;
  2. derives a ModelIdentity and asset fingerprint;
  3. decodes model-specific preprocessing, tokenizer, function-name, normalization, and configuration metadata; and
  4. exposes a SimilarityBackendDescriptor containing all compatibility-critical versions.

The descriptor is effectively the type identity of an embedding on disk. It records backend, model fingerprint, representation, preprocessing, normalization, and configuration versions. Image artifacts additionally record vector dimensions, schema version, and a source fingerprint. Consequently, RawCull does not compare a DataComp vector with an OpenAI vector, reuse an artifact after preprocessing changes, or silently treat an edited source file as unchanged.

Image preprocessing and inference

RawCullSimilarityImageDecoder first asks RawParserKitImageLoader for a bounded thumbnail. If that fails it attempts an ImageIO thumbnail without requesting a full fallback decode. The resulting CGImage is passed through the model-specific preprocessing declared by the CLIP bundle. Current metadata supports either the legacy stretch/bilinear path or shortest-side resize plus square center crop with bicubic interpolation and configured RGB mean/standard deviation.

The provider lazily loads Core AI functions and tokenizer resources. The image function receives the prepared image tensor and, if required by the exported graph, dummy text inputs. Its output is flattened, checked against the expected dimension, wrapped in an ImageEmbedding, JSON-encoded, and stored in a SimilarityArtifact. The backend descriptor records the bundle’s declared normalization version; image/text comparison later verifies that the image vector actually has approximately unit magnitude.

Indexing and recovery

RawCullCLIPSimilarityService uses PhotoAIKit’s bounded SimilarityArtifactIndexer with concurrency limit 1. CLIP inference is kept serial because the provider is actor-owned and model execution is resource intensive. There is no per-file or whole-batch Vision substitution during a CLIP pass.

For a non-finite vector, RawCullRecoveringCLIPArtifactProvider performs a targeted sequence:

  1. reject the invalid output;
  2. retry once with the already-loaded provider;
  3. construct a fresh provider from the same validated model location;
  4. verify that the replacement descriptor is exactly the same; and
  5. retry once with the replacement.

Successful CLIP artifacts from other files are retained. A file that still fails is reported and remains unindexed; it is not given a Vision artifact that would make the batch heterogeneous. Decode failures and inference failures are recorded separately for diagnostics.

Image similarity and burst grouping

For compatible normalized image embeddings, PhotoAIKit returns cosine distance. SimilarityScoringModel owns the artifact dictionary and computes distances from an anchor. It can apply a small RawCull-owned subject-label mismatch penalty, then uses those distances as input to ranking and burst grouping.

This responsibility split is deliberate:

  • CLIP defines the vector and mathematical comparison;
  • PhotoAIKit defines artifact compatibility and indexing mechanics; and
  • RawCull defines catalog admission, grouping thresholds, ordering, progress, persistence, and culling policy.

Vision remains the runtime fallback when CLIP is disabled or cannot be validated. Vision and CLIP artifacts are both descriptor-bearing, so switching backends causes incompatible state to be rejected or rehydrated rather than misinterpreted.

Semantic search never indexes missing images as a side effect. It operates only on already-persisted, descriptor-compatible CLIP image artifacts:

flowchart LR
    Query["Text query"] --> Tokens["CLIP tokenizer"]
    Tokens --> TextModel["CLIP text function"]
    TextModel --> TextVector["Validated normalized text vector"]
    Images["Compatible cached image artifacts"] --> Compare["Dot product / cosine similarity"]
    TextVector --> Compare
    Compare --> Sort["Descending score with deterministic tie breaks"]

RawCullCLIPSemanticSearchService trims and validates the query, filters image artifacts by the complete backend descriptor, generates one transient text embedding, and scores each compatible image. Because both vectors are normalized, their dot product is cosine similarity. Valid scores lie in -1...1; they are relative ranking values, not confidence percentages.

Sorting is deterministic: score descending, then original catalog order, localized filename, and UUID. Individual malformed artifacts become per-file failures rather than aborting every valid result. Text embeddings are scoped to one search and are not persisted.

Deep Analysis: How SAM 3 and CLIP Work Together

The UI calls this mode SAM 3 + CLIP, but the two models are not fused and CLIP does not calculate the final focus score. Their collaboration is a staged pipeline:

  1. CLIP optionally supplies a coarse subject label from existing embeddings.
  2. That label selects an ordered set of text prompts for SAM 3.
  3. SAM 3 creates or retrieves the best acceptable subject mask.
  4. RawCull measures detail only inside that mask and recommends the strongest candidate.

If CLIP semantic artifacts are unavailable, RawCull falls back to the existing saliency label from normal sharpness analysis. If neither label exists, SAM 3 still receives the general subject prompt. Therefore SAM 3 is the required model for Deep Review; CLIP enriches prompt selection when available.

1. Building the request

DeepAIReviewController.start(for:) asks RawCullViewModel for a stable group context. deepAIReviewContext(for:) captures:

  • a BurstGroupSignature tied to the current catalog and exact member files;
  • existing burst ranks, falling back to input order;
  • normal sharpness scores;
  • a subject label;
  • the normalized camera autofocus point; and
  • the chosen sharpness source: embedded preview or RAW demosaic.

For the CLIP label pass, SimilarityScoringModel.classifySubjects runs six literal queries against existing semantic artifacts: person, bird, deer, animal, car, and landscape. Each file receives the label with its highest cosine similarity. This pass does not alter the visible semantic-search result, decode source images, or generate missing embeddings.

2. Candidate limiting and decoding

Candidates are sorted by current burst rank. Groups of 12 or fewer are analyzed in full; larger groups analyze the first 8 candidates. Each candidate is then decoded at a bounded size. Embedded-preview mode reuses RawCull’s similarity decoder. RAW-demosaic mode uses CIRAWFilter, explicitly sets sharpness to 0, detail to 0.6, contrast to 1, and exposure to 0, then downsizes to the configured maximum before producing a CGImage.

3. Prompt selection

The preset and subject label determine ordered attempts:

Preset/evidenceSAM prompt order
Full Subjectsubject
Auto or Head/Face with bird/wildlife labelbird head, bird, subject
Auto or Head/Face with person/face labelface, person, subject
Auto or Head/Face with deer labelanimal head, deer, animal, subject
Auto or Head/Face with generic animal labelanimal head, animal, subject
No recognized labelsubject

The Head/Face preset is considered verified only when the first selected prompt is bird head, animal head, or face; falling back to a broader mask is reported as specificPromptNotFound.

4. SAM 3 inference

PhotoAIKit’s CoreAISAM3Provider is actor-isolated. It validates the bundle and lazily creates a CoreAISegmentationEngine plus CLIP-compatible text tokenizer. The prompt text is tokenized and sent with the bounded image to the Core AI segmenter. Runtime parameters use a 0.5 mask threshold and at most 5 segments.

The response’s probability map is preferred. If absent, the provider unions compatible returned segment masks. Probabilities are converted to a white RGBA mask with a smooth alpha transition around the threshold. The result records the prompt, confidence, model identity, input/output sizes, timing, resource, and asset identity.

PhotoAIKit’s SegmentationService first checks memory and disk stores by a key that includes source file identity, prompt, model identity, and maximum input size. A missing mask is generated, resized back to the display image dimensions, and saved to both stores. Model changes therefore do not accidentally reuse masks made by a different SAM asset.

SubjectMaskSelector tries prompts in order, measuring coverage and quality for each candidate. It stops at the first mask meeting the warning-or-better threshold, otherwise retains the best attempt by quality, confidence, and then coverage. Every attempt records cache miss, candidate quality/confidence, or failure.

5. Subject-detail scoring

SubjectMaskFocusScorer converts the photograph to luminance using Rec. 709 weights and samples Laplacian-style edge energy. Only pixels with mask alpha above 16 contribute to subject evidence. It computes:

  • broad subject score — robust tail detail across the masked subject;
  • local detail score — the strongest reliable cell in a 6 × 6 patch grid;
  • fine detail score — micro-contrast within the mask;
  • mask coverage — masked pixels divided by total pixels; and
  • AF evidence — whether the normalized autofocus point lies inside the mask.

The final detail score is:

0.40 × broad subject detail
+ 0.40 × strongest local detail (or broad detail when local is unavailable)
+ 0.20 × fine detail

If global edge detail exceeds the subject score by the configured margin, RawCull applies a 0.82 background-dominance multiplier and records a caution. The scorer reports missing local patches, unusable masks, unavailable subject detail, and other evidence limitations rather than manufacturing a score.

6. Recommendation and confidence

Candidates sort by deep score descending, with the earlier burst rank breaking ties. The first finite score becomes the recommendation. Reasons record strong subject detail, AF-inside-subject evidence, local detail evidence, and successful prompt matching.

Confidence depends on the winning margin and evidence quality:

  • high: at least a 12% lead, mask and local evidence present, no issues, and no fallback prompt;
  • medium: at least a 5% lead, or strong evidence obtained through a fallback prompt; and
  • low: all other cases.

The result is advisory. Deep Review stores results by group signature, publishes progress after every candidate, and keeps completed candidate/mask evidence for the analysis history and zoom outline. It does not silently change ratings or apply a culling decision.

How Qwen Analysis Works

Qwen is not part of CLIP similarity or SAM segmentation. It is a separate local vision-language tool selected in the AI Analysis view.

Validation and lazy loading

QwenInferenceRuntime is an actor with three pieces of state: an optional CoreAIQwenProvider, an optional loaded CoreAIVisionLanguageModel, and a generation counter. Validation asks PhotoAIKit’s Qwen factory to inspect the bundle. The provider checks required tokenizer resources, metadata kind, Qwen identity, vocabulary/context values, and VLM-specific embedding and vision assets. RawCull additionally rejects a valid text-only Qwen bundle because photo assessment requires modality .vision.

Successful validation stores the lightweight provider and clears any previously loaded model. The first call to assess constructs the heavy vision-language model asynchronously; later calls reuse it. A generation captured before loading prevents a model removed or replaced during the await from becoming active afterward. clear() advances the generation and releases both provider and loaded model.

Per-image request

The feature processes pending files sequentially. It requests a thumbnail up to 2048 pixels, then calls assess(criteria:image:). The runtime creates a new Foundation Models LanguageModelSession around the reused model, attaches the CGImage, and requests at most 512 response tokens.

The prompt includes the user’s criteria and asks for exactly one JSON object when the request is a photo assessment. The schema contains:

  • subject description;
  • composition, exposure, and subject-visibility scores from 1 through 5;
  • optional eyesOpen;
  • up to four problems and strengths; and
  • confidence from 0 through 1.

The instruction explicitly says to use visible evidence only. For a request that does not fit the assessment schema, Qwen may answer in ordinary text.

Response handling

QwenModelResponse.decode trims the response. It first attempts to extract the outermost JSON object and decode QwenPhotoAssessment. Score ranges and confidence are validated. If that structured decode does not succeed but the response is nonempty, RawCull retains it as a free-form result.

The structured overallScore is a RawCull presentation value, not a model output:

0.50 × composition
+ 0.20 × exposure
+ 0.30 × subject visibility

Each component is first divided by 5. The batch continues after an individual file fails, and successful or failed results replace earlier results for the same file. Already analyzed files are skipped on the next run. Cancellation and model removal advance operation state so obsolete work cannot publish as a current result.

Qwen results currently live in the feature’s in-memory results array; unlike CLIP artifacts and SAM masks, this implementation does not persist them across application sessions.

Objects: instance-level SAM 3 and Qwen analysis

Objects is a third AI Analysis tool beside SAM 3 + CLIP and standalone Qwen. It accepts selected Grid photos or tagged photos. It requires installed, validated SAM 3 and vision-capable Qwen. CLIP embeddings and Deep Review’s single-subject score do not feed this workflow.

End-to-end stages and ownership

  1. The stable, main-actor object feature checks that both model services are available and snapshots the concept mode and photographic criteria.
  2. It loads one bounded RAW or JPEG thumbnail, at most 4,320 pixels on its longest side, and processes files sequentially.
  3. Automatic mode asks Qwen for visible object concepts; Specific Concepts parses the user’s comma-separated noun phrases.
  4. PhotoAIKit’s object service asks SAM 3 for up to eight instances per concept and checks its separate object-mask caches.
  5. RawCull filters weak/invalid masks, merges near-duplicate regions across concepts, and assigns board-local IDs 1 through 8.
  6. A deterministic 2,048-pixel board shows the original overview above numbered, outlined object crops. Qwen assesses that one photograph.
  7. RawCull extracts one complete JSON object and validates the schema, every expected board ID, values, list caps, and finite confidence. The UI shows per-object findings or a visible, retryable assessment problem.

The availability state distinguishes checking, ready, SAM 3 unavailable, Qwen unavailable, and both unavailable. Settings installs the current segmentation service and Qwen status into the existing feature after validation. A changed service or status cancels active work. Switching AI tools or input source also cancels an active batch. Completion, result replacement, and cancellation are generation-gated.

Concept discovery and manual mode

Automatic asks Qwen for zero to six short concept entries. Each entry has query, displayName, and reason. Query must be a concrete, visible, whole-object noun phrase suitable for SAM 3. SegmentationConcept validates it; normalized duplicates collapse. An empty set or invalid JSON is an actionable discovery failure, with no guessed fallback concepts. The discovery request cap is 384 output tokens.

Specific Concepts bypasses discovery. The user enters up to six comma-separated queries such as bird, person. Empty or invalid entries fail before segmentation; duplicate normalized queries collapse. Both modes can add photographic criteria to the final Qwen request.

SAM 3 instances, filtering, and caches

ObjectSegmentationService uses source file identity, concept, SAM 3 model identity, the 4,320-pixel input limit, and the eight-instance limit in its cache key. It checks the object-mask memory and optional disk store before inference, bounds the image, invokes CoreAISAM3Provider.segmentInstances, resizes masks to display dimensions, and saves the typed result. This is independent of Deep Review’s SegmentationService and SubjectMaskSelector, which choose one subject mask from ordered prompt attempts.

RawCull counts the raw SAM 3 candidates. Its deduplicator rejects nonfinite or below-0.5 mask scores, invalid boxes, masks with fewer than 64 of 256-by-256 sampled pixels, and masks covering at least 95% of the sample. Candidates sort by score and geometry. Overlapping masks with similar area merge when mask intersection-over-union reaches 0.85 or smaller-mask containment reaches 0.90; a second concept becomes an alias of the retained object. At most eight objects remain. Their board IDs are strings and local to that analysis result. A SAM 3 mask score describes segmentation quality; it is distinct from Qwen’s assessment confidence.

An empty retained set is a successful No Matching Objects result and does not call Qwen for a board. Genuine no-match photos still need broader validation.

Board geometry and the Qwen contract

ObjectReviewBoardRenderer creates one 2,048-by-2,048 image. Its top half is an aspect-fit overview of the source photo. Its bottom half contains up to eight padded crops in two rows of four. Each crop and yellow mask outline are aspect-fit. A separate dark header carries a large white number without covering the subject. The renderer explicitly converts normalized bottom-left box coordinates to top-left CGImage crop coordinates.

If the source or mask crop for a retained object cannot be made, rendering throws reviewBoardUnavailable. The feature records the per-photo failure before asking Qwen to assess a board; it does not submit a board with a numbered ID whose crop was silently omitted.

The prompt tells Qwen that the overview and crops repeat views of one photograph, each board ID denotes a different physical subject, and the objects array must contain exactly one entry for each ID. Descriptions should use the matching numbered crop; relationships should use the overview. The assessment request cap is 1,024 output tokens. The requested result has an optional scene summary; per-object concept, description, visibility, focus, expression, obstructions, strengths, problems, and confidence; and photo-level relationships, strengths, problems, preferred IDs, and confidence.

ObjectJSONEnvelope extracts one complete balanced JSON object from a recoverable Markdown or prose wrapper, respecting quoted braces. It rejects incomplete JSON and multiple objects. The decoder accepts only the narrow variations observed from the packaged model: an omitted imageSummary, a whole-number board ID normalized to a string, and one short text value in an object list field (the exact string “none” becomes an empty list). It still requires every board ID exactly once, rejects unknown or duplicate IDs and preferred IDs, enforces list limits and enum values, and requires finite confidence in 0…1. Free-form prose is not upgraded to a structured judgment.

If Qwen fails this boundary, the detail view retains its response and shows a specific assessment error. The results table says Assessment needs retry. Retry reuses successful SAM 3 masks only when source size/date, Qwen model name, concept mode and queries, and cache keys still match. Otherwise it segments again. The table labels Qwen confidence; the detail view labels SAM 3 mask score separately and shows numbered boxes, cached-mask outlines, the selected crop, and its assessment.

The object count in the table is the number of retained SAM 3 matches for the chosen concepts, after filtering and deduplication. It is not a count of all subjects in the photograph. A Complete row means the response passed the schema and board-ID checks; the photographer still needs to compare its descriptions, count claims, and confidence with the source image. The detail panel shows the assessment for the currently selected object, not all object descriptions at once.

September 24, 2026 packaged-model check

Eight supplied puffin ARW files were processed in both modes with local qwen3_vl_2b and sam3_float16.aimodel. Automatic discovered puffin for all eight; manual mode used bird. All 16 runs produced structured assessments with the expected board ID set. SAM 3 returned seven or eight raw candidates per run, filtered to one retained bird on six photos and two birds on two. Warm concept discovery took 2.98–3.56 seconds, segmentation 6.59–6.99 seconds, board rendering 0.028–0.044 seconds, and final Qwen assessment 10.21–20.14 seconds. The sequential probe took 371.5 seconds; process peak resident memory was 9.47 GiB. The first Automatic run included model startup and took 45.1 seconds overall.

These measurements came from an in-process macOS feature/model test, not a release build, clean install, or TestFlight run. Separate two-bird visual checks found a swapped flying/perched description, nearly identical descriptions for differently facing birds, and a response claiming three birds where the photograph had two. Those errors occurred with schema-valid output and Qwen-reported confidence of 0.95–1.00. Valid JSON and IDs verify response shape, not visual accuracy or calibrated confidence. Mixed categories, touching/overlapping and tiny subjects, genuine no-match cases, RAW/JPEG parity, model removal, large tagged batches, and lifecycle checks remain open. The per-photo table and gate status are in the RawCull implementation notes.

September 24–26, 2026 in-app observations

In-app Automatic runs showed all eight puffin photos and all eight photos in a mixed-subject batch as Complete. The mixed batch included landscape, deer, muskox, horse, bird, rabbit, and puffin photographs. These screenshots show the workflow operating in the app on those selections, but do not identify the installed model-pack fingerprints or establish a clean-install TestFlight run.

Two visual checks remain unresolved. In _DSC3028.ARW, clicking the two numbered puffins showed crops and descriptions associated with opposite birds; the source of the mismatch, whether Qwen’s board grounding or the UI’s object-to-crop association, has not been established. In _DSC3031.ARW, SAM 3 retained two puffins and Qwen’s object details described two positions, while its image summary claimed a third puffin on the ground. That Complete result reported 95% Qwen confidence. The invented third bird is a prose error, not a third SAM 3 instance or a missing board ID.

Objects is being treated as an advisory test feature while users report issues. The two photographs are regression cases for visual grounding and object mapping. Before treating the broader validation as complete, the signed build and hosted model packs still need a clean-install run covering launch, analyze, cancel, remove, and reinstall, with build, model identities, macOS version, and memory recorded. See the in-app validation and follow-up and September 26 observation.

Private diagnostics

The feature records raw/retained counts, model identities, and separate concept-discovery, segmentation, board-rendering, and assessment durations. An explicit RAWCULL_OBJECT_CAPTURE_DIR environment variable enables private response capture with file name/ID, stage, model, requested token cap, and response character count. The directory must have 0700 permissions and new files use 0600. Normal logs do not include full images, prompts, or Qwen responses. The runtime exposes neither finish reason nor generated-token count, so response length and visible truncation must be inspected directly. Remove private capture files after diagnosis.

Capability and Failure Behavior

The settings UI distinguishes these states instead of reducing them to one Boolean:

  • checking: a location exists or is being resolved, but validation is not complete;
  • available: validation and provider construction succeeded;
  • missing: expected resources are absent;
  • invalid: a resource exists but metadata, checksum, files, modality, or provider construction failed; and
  • unavailable: a feature cannot be offered for another explicit reason.

Semantic-search readiness is separate from generic CLIP readiness. Image similarity can always fall back to Vision; text search requires a validated provider that implements the text/image contracts. Deep Review can keep its stable controller while its service is temporarily nil. Qwen publishes its own status and cancels an active batch if that status becomes unavailable.

Persistence Boundaries

DataLifetime/locationCompatibility protection
CLIP or Vision similarity artifactsPer-file analysis artifact store and burst cacheFull backend descriptor plus source fingerprint and schema version
CLIP text query embeddingOne search callNever persisted
SAM 3 masksMemory store plus Caches/no.blogspot.RawCull/SAM3Masks when disk-store construction succeedsSource identity, prompt, model identity, and max input side
Deep Review recommendationsIn-memory feature dictionary keyed by BurstGroupSignatureExact group signature; reset/cancellation generation
Qwen resultsIn-memory feature array keyed by file UUIDCurrent batch generation; no cross-launch persistence
Objects masksSeparate object-mask memory store and optional ObjectMaskDiskStoreSource identity, concept, SAM 3 model identity, 4,320-pixel input limit, and eight-instance limit
Objects assessments and timingsIn-memory feature results keyed by file UUIDBatch generation and board-ID validation; no cross-launch assessment persistence
User model selectionsUserDefaultsInclusion lists sanitize choices no longer shipped
Model assetsManaged Background Assets locationsCatalog ID, model bundle validation, and asset fingerprint/checksum

Practical Trace Points

When debugging a model problem, follow the layer that owns the decision:

  1. Asset not present or licence blocked: model download catalog, downloads model, and download service.
  2. Bundle present but invalid: RawCullAIModelResourceManager and PhotoAIKit ModelBundleResolver.
  3. Provider validates but feature stays on Vision: settings snapshot, RawCullAIModelRuntime.similarityService, and runtime configuration identity.
  4. Some CLIP images fail: decoder/inference failure report and finite-vector recovery in RawCullCLIPSimilarityService.
  5. Semantic search has no candidates: semantic artifact hydration and exact descriptor compatibility.
  6. SAM mask is missing or poor: prompt attempts, mask cache key, geometry, quality, and segmentation diagnostics.
  7. Deep score looks unexpected: inspect broad/local/fine evidence, mask coverage, AF inclusion, and background-dominance caution.
  8. Qwen is available but a batch fails: distinguish thumbnail decoding, lazy model load, session response, empty response, and per-file result decode.
  9. Objects fails before SAM 3: inspect concept discovery and its exact JSON or concept-validation error; Specific Concepts isolates that boundary.
  10. Objects has masks but no structured judgment: inspect the assessment error, rendered board, and board-ID set. Retry may reuse cached masks.
  11. Objects text disagrees with the photo: compare the source, numbered crop, outline, and description. High model-reported confidence does not settle a grounding error.

The central architectural rule is that model runtimes create typed evidence; RawCull’s feature and policy layers decide how that evidence affects ranking, grouping, presentation, and user actions.

4 - The RawCull AI Runtime

How RawCull owns and refreshes local CLIP, SAM 3, Qwen, Vision, Deep Review, and Objects runtimes.

The RawCull AI Runtime

RawCull’s AI runtime is the long-lived object graph that connects downloaded model assets to stable application features. It is not one model, one thread, or a background daemon. It is a set of objects with deliberately different lifetimes and actor-isolation rules.

The current implementation has two runtime layers:

RuntimePrimary responsibility
RawCullAIModelRuntimeOwn concrete provider/resource lifecycles: CLIP, SAM 3, Qwen, Vision, model capability snapshots, separate subject/object mask stores, and segmentation-service installation.
RawCullIntelligenceRuntimeOwn stable similarity, semantic search, Deep Review, Qwen, and Objects feature lifetimes; apply complete, revisioned similarity/semantic/segmentation settings decisions without rebuilding the graph.

That distinction replaces the older, broader RawCullAIIntegration shape. The authoritative sources are RawCullAIModelRuntime.swift and RawCullIntelligenceRuntime.swift. For model algorithms and data products, see AI Models in RawCull.

Runtime Topology

flowchart TD
    App["RawCullApp"] --> AppState["RawCullApplicationState"]
    AppState --> VM["RawCullViewModel"]
    AppState --> Runtime["RawCullIntelligenceRuntime"]

    Runtime --> ModelRuntime["RawCullAIModelRuntime"]
    Runtime --> Settings["RawCullAISettingsModel"]
    Runtime --> Downloads["RawCullAIModelDownloadsModel"]
    Runtime --> Similarity["RawCullSimilarityFeature"]
    Runtime --> Semantic["RawCullSemanticSearchFeature"]
    Runtime --> Review["DeepAIReviewController"]
    Runtime --> Qwen["RawCullQwenAnalysisFeature"]
    Runtime --> Objects["RawCullObjectAnalysisFeature"]

    ModelRuntime --> CLIP["CLIP resource managers/providers"]
    ModelRuntime --> SAM["SAM 3 resource manager/provider"]
    ModelRuntime --> QwenActor["QwenInferenceRuntime actor"]
    ModelRuntime --> Vision["Vision fallback"]
    ModelRuntime --> Masks["Mask repository/stores/selector"]
    ModelRuntime --> ObjectMasks["Separate object-mask stores and instance service"]

    Similarity --> SharedModel["SimilarityScoringModel"]
    Semantic --> SharedModel
    Review --> DeepFeature["DeepAIReviewFeature"]
    Qwen --> QwenActor
    Objects --> QwenActor
    Objects --> ObjectMasks

The arrows above mix ownership and collaboration. The precise ownership rules are discussed below; notably, callbacks from a child toward an owner are weak.

What Each Layer Owns

RawCullAIModelRuntime

The model runtime is @MainActor because model selection and provider installation must be coordinated with observable feature state. Heavy work is still isolated elsewhere: its resource managers and Qwen runtime are actors, and PhotoAIKit’s CLIP and SAM 3 providers are actor-owned.

It owns:

  • application AI paths;
  • three RawCullAIModelResourceManager actors: SAM 3, DataComp CLIP, and OpenAI CLIP;
  • a single QwenInferenceServing actor;
  • the always-available VisionFeaturePrintBackend and Vision similarity service;
  • dictionaries of validated CLIP and segmentation providers;
  • resolved CLIP model locations used to create replacement providers;
  • memory and optional disk subject-mask stores;
  • separate memory and optional disk object-mask stores;
  • an optional ObjectSegmentationService built from the validated SAM 3 provider and those object stores;
  • the current SubjectMaskRepository, SegmentationService, and SubjectMaskSelector;
  • the selected segmentation model and active model identity; and
  • the latest RawCullAICapabilities snapshot.

Views do not traverse this object. They receive the focused feature surfaces from RawCullIntelligenceRuntime.

RawCullIntelligenceRuntime

The intelligence runtime owns the objects whose identities must remain stable:

let modelRuntime: RawCullAIModelRuntime
let similarityFeature: RawCullSimilarityFeature
let semanticSearchFeature: RawCullSemanticSearchFeature
let deepAIReviewController: DeepAIReviewController
let qwenAnalysisFeature: RawCullQwenAnalysisFeature
let objectAnalysisFeature: RawCullObjectAnalysisFeature
let settingsModel: RawCullAISettingsModel
let modelDownloadsModel: RawCullAIModelDownloadsModel

It also records the last accepted configuration revision and identity. Its single mutation entry point is apply(configuration:).

RawCullViewModel

The main view model owns product and catalog policy, not model runtimes. It answers questions such as which files are selected, what the current catalog identity is, what burst ranks and sharpness scores exist, and whether an AI operation conflicts with other work. Narrow protocols expose only the pieces the AI features need.

Construction: From App Launch to a Live Graph

RawCullApp.init() calls RawCullApplicationState.live(). SwiftUI retains the returned view model and intelligence runtime in separate @State properties:

RawCullApp
 ├─ strong → RawCullViewModel
 └─ strong → RawCullIntelligenceRuntime

live() constructs a default RawCullAIModelRuntime, then delegates to the injectable RawCullApplicationState.make(...). The factory accepts stores, preferences, scanners, download catalog/coordinator, and version information as parameters so tests can build the same graph with deterministic substitutes.

Phase 1: initialize the model runtime

RawCullAIModelRuntime.init performs only synchronous, bounded setup:

  1. Store RawCullAIPaths and the Qwen inference actor.
  2. Create SAM 3 and CLIP resource-manager actors with their PhotoAIKit factories. Managed URLs are initially unset.
  3. Create one Vision provider and wrap it in RawCullVisionSimilarityService.
  4. Create SubjectMaskMemoryStore.
  5. Attempt to create SubjectMaskDiskStore at Caches/no.blogspot.RawCull/SAM3Masks. Failure is represented as a capability state; it does not prevent the application from launching. Create a separate object-mask memory store and try an object disk store at the configured object-mask cache path. A disk-store failure leaves Objects with its memory store.
  6. Build the list of usable stores: memory always, disk when construction succeeded.
  7. Create an UnavailableSegmentationProvider, repository, segmentation service, and selector. This placeholder gives the graph a complete shape before SAM validation.
  8. Publish an initial capability snapshot: Vision available, model resources checking, and mask-storage status known.

No CLIP, SAM 3, or Qwen model engine is loaded in this initializer.

Phase 2: assemble stable features

RawCullApplicationState.make then performs the following order:

  1. Create RawCullAIModelDownloadsModel from runtime paths and the production model catalog.
  2. Create RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature with the same modelRuntime.qwenInference. Pass the object feature the exact memory and optional disk object-mask stores owned by the model runtime.
  3. Create DeepAIReviewFeature with the initial mask-generation capability.
  4. Bind that exact feature to the model runtime. The runtime installs an actual pipeline later when a segmentation provider becomes available.
  5. Create RawCullAISettingsModel with model runtime, downloads model, Qwen feature, user defaults, and saved-burst-evidence scan.
  6. Ask settings for a synchronous revision-0 configuration.
  7. Create one SimilarityScoringModel from the selected similarity service, semantic capability/service, and persistent artifact store.
  8. Wrap it in RawCullSimilarityFeature and RawCullSemanticSearchFeature. Both wrappers refer to the same scoring model.
  9. Wrap the Deep Review feature in DeepAIReviewController.
  10. Create RawCullViewModel with those exact feature/controller instances.
  11. Create RawCullIntelligenceRuntime and bind the similarity feature’s weak application context.
  12. Bind semantic search to the view model and settings to the runtime.

The final settings binding immediately publishes the first configuration. This happens only after every receiver exists.

Phase 3: verify graph identity

Debug assertions verify that:

  • view model and runtime share the same similarity feature;
  • semantic search and similarity share the same scoring model/feature identity;
  • view model and runtime share the same semantic and Deep Review objects;
  • controller and model runtime refer to the same Deep Review feature;
  • runtime and Qwen feature share the same Qwen inference actor;
  • runtime and Objects feature share that same Qwen inference actor;
  • settings, runtime, and Qwen feature share the intended model runtime and inference actor; and
  • settings and runtime expose the same downloads model.

These are architectural invariants. Two equivalent-looking instances would not share task handles, progress, caches, result dictionaries, generation counters, or SwiftUI observation.

Development Handoff: PhotoAIKit Objects to the Runtime

PhotoAIKit supplies provider factories, typed contracts, workflows, and stores. RawCull creates those objects and decides when they become usable. There is no package callback that injects providers into RawCullIntelligenceRuntime: RawCullAIModelRuntime owns the provider handoff, while Settings delivers selected services to the stable feature objects.

Declare and construct the package boundary

RawCullAIModelRuntime.swift imports CoreAICLIPBackend, CoreAISAM3Backend, PhotoAIContracts, PhotoAIStorage, PhotoAIWorkflows, and VisionFeaturePrintBackend. At construction it passes CoreAICLIPProvider.factory and CoreAISAM3Provider.factory to separate RawCullAIModelResourceManager actors. It also creates the Vision provider, SubjectMaskMemoryStore, optional SubjectMaskDiskStore, and an unavailable segmentation provider. The latter lets SegmentationService, SubjectMaskRepository, and SubjectMaskSelector exist before a SAM bundle is validated. Qwen’s CoreAIQwenProvider.factory is used by the separate QwenInferenceRuntime actor.

These are the objects crossing the package boundary:

PhotoAIKit object or contractRawCull receiver and use
CoreAICLIPProvider and SimilarityBackendDescriptorModel runtime retains the validated provider; similarity and semantic services wrap it. The descriptor identifies compatible artifacts.
CoreAISAM3Provider as SubjectSegmenting, with ModelIdentityModel runtime installs it in a segmentation service and selector, then gives Deep Review a pipeline. Identity determines when the pipeline must be rebuilt.
The same CoreAISAM3Provider as ObjectInstanceSegmentingModel runtime builds a separate ObjectSegmentationService with object-mask stores and gives it to the existing Objects feature through Settings.
VisionFeaturePrintBackendModel runtime wraps it in the always-ready Vision similarity service.
SubjectMaskMemoryStore and optional SubjectMaskDiskStoreRepository and segmentation service share the stores; the disk store also supports Deep Review mask loading.
ObjectMaskMemoryStore and optional ObjectMaskDiskStoreA separate instance-mask cache namespace, shared by object segmentation, detail outline loading, and assessment retry.
CoreAIQwenProviderQwen inference actor retains the validated provider and loads its vision-language model on first use; the Qwen feature holds that same actor.

Enable installed models and hand off providers

The development path from a model download to a live feature is:

Background Assets snapshot (complete model-ID → URL map)
  → Downloads model → Settings.applyManagedModelLocations
  → RawCullAIModelRuntime.applyManagedModelLocations
  → resource-manager actors / QwenInferenceRuntime
  → PhotoAIKit capability check and provider construction
  → RawCullAIModelRuntime.refreshCapabilities
  → Settings.configurationSnapshot → RawCullIntelligenceRuntime.apply
  → existing feature objects

The model runtime supplies managed URLs to each CLIP and SAM resource manager. Each actor checks a metadata snapshot, asks its PhotoAIKit factory to validate the candidate bundle, and constructs a provider only for an available resource. refreshCapabilities() loads those actors concurrently, then stores the validated providers, their resolved CLIP URLs, and a capability snapshot on the main actor. If SAM 3 validated, it also constructs the object instance service with the separate object-mask stores. Qwen instead validates or clears its actor in applyManagedModelLocations; Settings passes its status to the already-created Qwen feature. Validation does not eagerly load Qwen’s vision-language model.

The provider reaches a feature through one of three paths:

  1. For image similarity, Settings calls modelRuntime.similarityService(prefersCLIP:clipModel:). It wraps the selected validated CLIP provider in RawCullCLIPSimilarityService, or uses the Vision service. The CLIP service also receives the resolved bundle URL through a replacement-provider factory for finite-vector recovery.
  2. For semantic search, Settings calls modelRuntime.semanticSearchService(clipModel:). The same validated CLIP provider backs RawCullCLIPSemanticSearchService; a missing provider yields no semantic service. configurationSnapshot carries the selected capability and service to RawCullIntelligenceRuntime.apply(configuration:), which updates the shared SimilarityScoringModel through its stable feature.
  3. For Deep Review, refreshCapabilities() activates the selected SAM provider. The model runtime rebuilds the repository, segmentation service, and selector only when ModelIdentity changes, then calls DeepAIReviewFeature.install with a pipeline, optional disk-mask loader, and availability. The existing controller and feature keep their identities.
  4. For Objects, Settings installs modelRuntime.objectSegmentation and the current Qwen status into RawCullObjectAnalysisFeature after validation. The feature stays at the same identity and becomes ready only when both services are available. It uses the same Qwen actor as standalone Qwen and the same SAM 3 provider as Deep Review, through a different workflow/cache.

The runtime configuration carries selected services and descriptor-based identity, not an unvalidated model URL. Its revision prevents an older Settings decision from overwriting a newer one. When adding a PhotoAIKit backend, wire its factory and managed location into the model runtime, translate its capability and provider result, then expose it through the appropriate stable feature or configuration path. Keep model-specific inference inside the provider or actor and application policy inside RawCull.

Startup Refresh and Installed-Model Activation

The graph is usable immediately with Vision while disk checks happen later. When RawCullMainView appears, this task starts the real refresh:

.task {
    await intelligenceRuntime.settingsModel.refresh()
}

The call path is:

sequenceDiagram
    participant View as RawCullMainView
    participant Settings as RawCullAISettingsModel
    participant Downloads as RawCullAIModelDownloadsModel
    participant Models as RawCullAIModelRuntime
    participant Resource as Resource-manager actors
    participant Runtime as RawCullIntelligenceRuntime

    View->>Settings: refresh()
    Settings->>Downloads: refresh()
    Downloads-->>Settings: applyManagedModelLocations(snapshot)
    Settings->>Models: applyManagedModelLocations(snapshot)
    Models->>Models: validate or clear Qwen
    Models-->>Settings: Qwen status
    par model validation
        Settings->>Models: refreshCapabilities()
        Models->>Resource: load SAM 3 and both CLIP choices
    and saved evidence
        Settings->>Settings: scan burst caches
    end
    Models-->>Settings: capabilities
    Settings->>Runtime: apply(revisioned configuration)
    Runtime-->>Settings: current capabilities

One complete location snapshot

applyManagedModelLocations(_:) is the only activation path for a complete set of installed locations. It gives the current SAM and CLIP URLs to their resource managers. A missing Qwen URL calls qwenInference.clear(); a present URL is standardized and validated.

Using a complete snapshot avoids a transient mixture such as “new CLIP, old SAM, removed Qwen still active.” Every invocation describes one model-install state.

Generation-gated refresh

Settings increments refreshGeneration before beginning work. The generation is checked after Qwen validation and after the concurrent capability/evidence work. A later refresh therefore supersedes an earlier one; the earlier result cannot publish merely because its disk work finished last.

The defer that clears isScanningSavedBurstData also checks the generation, so an obsolete refresh cannot hide the current refresh’s progress indicator.

Concurrent capability validation

RawCullAIModelRuntime.refreshCapabilities() starts SAM 3, DataComp CLIP, and OpenAI CLIP loads with async let. Each resource-manager actor computes a lightweight recursive metadata snapshot. When unchanged, the previous validated capability/provider result is reused. When changed, PhotoAIKit validates the bundle and constructs a provider.

After all three complete, the main-actor runtime atomically replaces its provider dictionaries and resolved location dictionaries, translates package statuses into RawCull statuses, builds semantic-search readiness separately, stores one new capability snapshot, and activates the selected segmentation provider.

Capability State Is More Than “Loaded”

RawCullAICapabilityStatus distinguishes:

StateMeaning
checking(expectedLocations:)Validation is pending.
available(location:)Bundle validation and provider construction succeeded, or an always-available service such as Vision is ready.
missing(expectedLocations:)No candidate bundle was found.
invalid(location:reason:)A candidate exists but validation or provider construction failed.
unavailable(reason:)The runtime cannot offer the capability for another explicit reason.

CLIP model availability and semantic-search readiness are separate. A provider must expose the text/image contracts before semantic search is .ready. Vision remains a valid similarity service but can never satisfy semantic text search.

Qwen uses an internal QwenModelStatus with not-configured, checking, available, missing, and invalid states. Settings translates it to the common capability presentation and updates the Qwen feature at the same time.

Resource Managers and Their Cache

RawCullAIModelResourceManager<Provider> is an actor because filesystem inspection, cryptographic bundle validation, and provider initialization must not run on the main actor or race with a managed-location change.

The resource cache has two keys:

  • RawCullAIModelResourceSnapshot, a sorted list of path, file kind, byte count, modification time, and symlink target; and
  • the resulting capability, optional provider, and optional provider-init failure.

setManagedCandidateURL invalidates both when the URL changes. load() also detects modifications within the same directory. The snapshot only decides whether validation may be reused. PhotoAIKit’s resolver remains responsible for metadata, required files, asset extension, fingerprint, and checksum validity.

Bundle validity and provider construction are reported separately. A bundle can be structurally valid yet fail to initialize its concrete runtime; RawCull maps that case to .invalid with the provider error so Settings can explain the actual stage that failed.

Selecting the Similarity Runtime

RawCullAIModelRuntime.similarityService(prefersCLIP:clipModel:) has a strict selection order:

  1. If the user disabled CLIP, return the existing Vision service.
  2. If the selected CLIP provider is absent, log the expected/resolved path and return Vision.
  3. If the provider exists but its resolved location is missing, return Vision.
  4. Otherwise return a new RawCullCLIPSimilarityService around the validated provider and supply a factory that can reconstruct a provider from the exact validated location for finite-vector recovery.

The service value can change while RawCullSimilarityFeature and SimilarityScoringModel retain their identities. Backend descriptors determine whether existing artifacts are still compatible.

Semantic search is constructed only from a currently validated CLIP provider. The provider itself satisfies both TextEmbeddingProviding and ImageTextSimilarityComparing, so RawCullCLIPSemanticSearchService can use the same model identity as the cached image embeddings.

Installing and Replacing the Segmentation Runtime

The model runtime retains the selected RawCullSegmentationModel, currently SAM 3. Selection and availability changes converge on activateSelectedSegmentationProvider(availability:).

When a provider is available, installSegmentationProviderIfNeeded compares its ModelIdentity with the active identity. A change rebuilds:

  1. SubjectMaskRepositoryConfiguration with prompt, model identity, and maximum input side;
  2. SubjectMaskRepository over the retained stores;
  3. SegmentationService over the new provider and stores; and
  4. SubjectMaskSelector over the matching repository and service.

When unavailable, the runtime installs the placeholder provider only if a real identity was previously active. Identity checks prevent needless reconstruction on repeated equivalent refreshes.

The stable DeepAIReviewFeature then receives:

  • a new RawCullDeepAIReviewPipeline when the provider exists;
  • a disk-mask loader when disk storage exists; and
  • the current availability state.

If availability disappears during a review, DeepAIReviewFeature.install cancels the active operation. Stored feature identity and already completed results remain under one owner.

Objects Runtime: shared models, separate workflow

Objects uses the same validated SAM 3 provider as Deep Review, but calls its instance-segmentation contract through a separate ObjectSegmentationService. It uses the same QwenInferenceRuntime actor as standalone Qwen, but owns a separate feature state machine and result array. These shared actors prevent duplicate heavy model instances; the separate workflows keep subject-mask selection and per-object instance analysis from sharing incompatible caches.

Activation and deactivation

The model runtime creates ObjectMaskMemoryStore at launch and tries ObjectMaskDiskStore using the configured object-mask cache directory. It starts with no object segmentation service. A complete managed-location snapshot clears that service, changes SAM 3 and CLIP resource-manager candidates, and validates or clears Qwen. Settings first cancels and marks Objects unavailable while this snapshot is applied. The generation-gated capability refresh then loads the SAM 3 provider; when available, it builds ObjectSegmentationService with the object stores, 4,320-pixel maximum side, and eight-instance maximum. Settings installs that service and the latest Qwen status into the existing RawCullObjectAnalysisFeature.

The feature is ready only when both services exist. It distinguishes checking, SAM 3 unavailable, Qwen unavailable, and both unavailable so the view can direct users to the missing download. install(segmentation:qwenStatus:) cancels an active batch when either dependency changes. Model removal therefore cannot leave an active Objects operation attached to a stale provider.

Feature lifetime and task boundaries

RawCullApplicationState.make creates one object feature and passes it to Settings, RawCullIntelligenceRuntime, and the AI Analysis view. Identity assertions check that it shares the model runtime’s Qwen actor. The feature retains the Qwen actor, image loader, mask stores, current object service, availability, results, progress, and a cancellable task. It snapshots the chosen concept mode, manual phrases, and criteria at batch start. Files are processed sequentially, each concept is segmented sequentially, and cancellation is checked between image loading, discovery, segmentation, board construction, and Qwen assessment.

The model runtime is main-actor isolated for installation and selection. ObjectSegmentationService and QwenInferenceRuntime are actors. The mask deduplication and board rendering CPU work runs in concurrent tasks before returning immutable results to the main-actor feature. Object results and captured timings stay in memory; object masks may outlive a feature result in the optional disk store. A detail view retrieves masks using the stored SAM 3 model identity and the same source/concept/cache parameters, then generates yellow outlines without regenerating a model result.

Board rendering must produce a crop and matching mask crop for every retained object. If either crop cannot be prepared, ObjectReviewBoardRenderer throws reviewBoardUnavailable; the feature records a per-photo failure before Qwen is asked to interpret the board. This keeps the board’s numbered panels and the set of IDs requested from Qwen in agreement.

No-match, response failure, and retry

An empty retained instance set completes as No Matching Objects and skips the Qwen board. When SAM 3 found instances but Qwen returned invalid or incomplete assessment JSON, the result keeps the instances and free-form response, records a stage-specific error, and remains eligible for Analyze or Retry Failed. It does not become Complete merely because segmentation succeeded.

For assessment retry, the feature checks source size and modification date, discovery mode, manual concepts, Qwen model name, and the stored result. It loads every retained mask using PhotoAIKit’s object cache key, which also encodes source identity, concept, SAM 3 model identity, input maximum side, and instance limit. Only a complete cache hit reuses segmentation; otherwise the workflow rediscovers concepts when needed and reruns SAM 3. Cancellation and generation checks keep a superseded result from publishing.

The results table labels whole-photo Qwen confidence. The detail view labels each SAM 3 mask score independently. The exact board-ID check establishes that the response describes the expected number of objects; it cannot prove that the descriptions correctly match those objects. The displayed object count is the number of retained SAM 3 matches for the requested concepts, not a census of the whole photograph. The September 24 puffin evaluation produced 16/16 structured results yet still exposed a swapped two-bird description and false scene claims. In later in-app checks, _DSC3028.ARW showed opposing crop and description associations for two birds, with the source of the mismatch still unresolved. _DSC3031.ARW retained two birds but its Complete, 95%-confidence Qwen summary invented a third. The detail panel displays only the currently selected object’s assessment. See AI Models in RawCull for the prompt, mask filtering, board layout, timings, in-app observations, and remaining validation work.

Qwen Runtime Lifetime

Qwen intentionally does not use the generic resource-manager actor. Its QwenInferenceRuntime owns two lazy layers:

  1. CoreAIQwenProvider, created during validation; and
  2. CoreAIVisionLanguageModel, created on the first assessment and retained for subsequent sessions.

Every validation or clear increments modelGeneration. assess captures that generation before an asynchronous model load and checks it after every suspension. If Settings removes or replaces the model during loading or generation, the old operation throws cancellation instead of publishing through an obsolete model.

RawCullQwenAnalysisFeature and RawCullObjectAnalysisFeature each own a separate batch generation and task while sharing that one inference actor. Both process images one at a time, isolate per-file failures, and cancel when their required model status changes. The features and inference actor protect different races: batch/UI lifetime and provider/model lifetime.

Revisioned Configuration Application

Settings publishes one complete RawCullIntelligenceConfiguration containing:

  • a monotonically increasing revision;
  • the selected similarity service;
  • semantic-search capability and optional service; and
  • the selected segmentation model.

Its identity contains values, not provider references:

  • selected similarity backend descriptor;
  • accepted artifact backend descriptors;
  • semantic capability;
  • semantic backend descriptor; and
  • segmentation selection.

Concrete services stay on @MainActor; only descriptor-based identity is Sendable.

The apply algorithm

RawCullIntelligenceRuntime.apply(configuration:) follows this order:

  1. Compute incoming identity.
  2. Reject a revision less than or equal to the last accepted revision. A same-revision/different-identity assertion catches a broken publisher.
  3. If a newer revision describes the same identity, record the newer revision without resetting any work.
  4. If segmentation selection changed, ask the model runtime to activate it.
  5. If similarity backend or accepted artifact descriptors changed, replace the similarity service through the stable feature.
  6. If semantic capability or semantic backend changed, replace semantic configuration through that same stable similarity feature.
  7. Record the accepted identity and revision.
  8. Return the model runtime’s current capability snapshot to Settings.
sequenceDiagram
    participant UI as Settings UI
    participant Settings as RawCullAISettingsModel
    participant Runtime as RawCullIntelligenceRuntime
    participant Models as RawCullAIModelRuntime
    participant Feature as RawCullSimilarityFeature

    UI->>Settings: change CLIP/model/segmenter preference
    Settings->>Settings: persist and increment revision
    Settings->>Runtime: apply(complete snapshot)
    Runtime->>Runtime: reject stale revisions and compare identity
    opt segmentation changed
        Runtime->>Models: setSelectedSegmentationModel
    end
    opt similarity changed
        Runtime->>Feature: replaceSimilarityService
    end
    opt semantic configuration changed
        Runtime->>Feature: replaceSemanticSearchConfiguration
    end
    Runtime-->>Settings: current capabilities

Revision and identity solve different problems. Revision orders decisions; identity determines whether a newer decision requires work.

What Service Replacement Invalidates

RawCullSimilarityFeature.replaceSimilarityService first compares complete backend identities. For a real change it:

  1. asks the application context to cancel and reset burst analysis tied to the old backend;
  2. installs the new service in SimilarityScoringModel;
  3. cancels existing image hydration;
  4. advances the image-hydration generation; and
  5. rehydrates the current catalog for the new accepted descriptors.

Semantic replacement updates semantic capability/service, cancels semantic hydration, advances its independent generation, and rehydrates compatible artifacts. Image similarity and semantic search use separate task handles and generations so changing one concern does not confuse completion from the other.

Catalog hydration performs image and semantic hydration in order and finally checks the current catalog identity. Ranking captures operation generation, catalog identity, and backend identity. A late completion must match all three before it is accepted.

Stable Identity: Why the Graph Is Not Rebuilt

Several views and owners retain the same feature objects:

RawCullApp ───────────→ RawCullIntelligenceRuntime
RawCullViewModel ─────→ RawCullSimilarityFeature A
Runtime ──────────────→ RawCullSimilarityFeature A
SwiftUI view ─────────→ RawCullSimilarityFeature A

On a model switch RawCull keeps A and changes its service. Rebuilding would create a split graph where an existing view observes A while runtime commands reach B. The objects might have the same type, but they would not share:

  • active task handles;
  • cancellation and generation state;
  • indexing/search progress;
  • hydrated artifacts and distances;
  • Deep Review results and mask-candidate history;
  • Qwen results and current batch;
  • Objects results, numbered instance IDs, retry state, and mask-cache access;
  • SwiftUI observation registrations; or
  • application-context bindings.

Keeping the stateful owner stable also lets the owner make a precise invalidation decision. Reconstructing everything would either lose unrelated state or risk copying backend-specific state into an incompatible runtime.

Ownership and Weak Coordination Edges

The runtime strongly owns Settings, but Settings must call back to the runtime. That callback is weak:

@ObservationIgnored private weak var configurationConsumer:
    (any RawCullIntelligenceConfigurationApplying)?

AnyObject permits weak protocol storage. @ObservationIgnored prevents a coordination detail from becoming UI state; it does not affect ownership.

Other coordination edges follow the same rule:

Strong owner/child relationWeak or per-run callback
Intelligence runtime → SettingsSettings → RawCullIntelligenceConfigurationApplying
Settings → Downloads modelDownloads model → RawCullAIManagedModelLocationsApplying
Runtime/view model → Similarity featureSimilarity feature → RawCullSimilarityApplicationContext
Runtime/view model → Semantic featureSemantic feature → RawCullSemanticSearchApplicationTarget
Runtime/view model → Deep Review controllerController → DeepAIReviewApplicationContext
View model → Burst coordinatorPer-run closures capture [weak self]

This produces one clear ownership direction and prevents retain cycles in both the full application session and shorter-lived tests.

Actor Isolation and Work Placement

ComponentIsolationWhy
RawCullAIModelRuntime@MainActorAtomically publishes capabilities and installs services used by observable features.
RawCullIntelligenceRuntime@MainActorApplies ordered settings decisions to stable UI-facing objects.
Settings/features/controllers/scoring model@MainActorOwn observable state, tasks, progress, and presentation.
RawCullAIModelResourceManageractorSerializes location/cache state while keeping filesystem and provider setup off the main actor.
CoreAICLIPProvideractorOwns lazy Core AI model/tokenizer state and serial inference.
CoreAISAM3ProvideractorOwns lazy segmentation engine/tokenizer state.
SegmentationServiceactorCoordinates provider access and mask stores.
ObjectSegmentationServiceactorCoordinates per-concept SAM 3 instance inference and the separate object-mask stores.
QwenInferenceRuntimeactorOwns provider, loaded VLM, and model generation.
RawCullObjectAnalysisFeature@MainActorOwns availability, sequential batch, progress, results, retries, and generation.
Pure scoring/ranking functions@concurrent or nonisolatedRun CPU-heavy work without making observable state unsafe.

The main actor coordinates; it does not perform model hashing, model execution, image vector comparison, or pixel-level subject-detail scoring itself.

Cancellation and Stale-Result Defences

RawCull uses several independent tokens because they protect different scopes:

Counter or identityRejects
Settings refreshGenerationAn older model/evidence refresh finishing after a newer one.
Runtime configuration revisionAn older settings decision arriving after a newer decision.
Similarity hydration generationsResults from tasks invalidated by service or catalog changes.
Similarity ranking generation + catalog/backend identityRanking for an old anchor, catalog, or backend.
Deep Review generationProgress/results after cancellation or restart.
Qwen feature generationBatch results after cancellation/restart.
Objects feature generationObject results after cancellation, model replacement, tool/source switch, or restart.
Qwen model generationA lazy load or response using a removed/replaced provider.

Task cancellation is cooperative, so the generation and identity checks are essential. Cancellation requests work to stop; generations prevent late work that did not stop immediately from becoming current state.

Failure and Fallback Policy

Runtime fallback is explicit:

  • If CLIP is disabled, missing, invalid, or fails provider construction, similarity uses Vision.
  • Once a CLIP indexing pass begins, individual failures do not receive Vision artifacts. Valid CLIP results remain; failed files stay unavailable.
  • Semantic search is unavailable without a compatible text-capable CLIP provider and compatible cached image artifacts.
  • Deep Review is unavailable without the selected segmentation provider. A failed candidate does not prevent later candidates from being evaluated.
  • Disk mask-cache creation failure leaves memory storage usable, while the capability reports the disk failure.
  • Qwen absence clears its runtime; an invalid or text-only bundle is reported, and an active batch is cancelled when status becomes unavailable.
  • Objects requires both SAM 3 and Qwen. Its availability names the missing dependency; it does not silently substitute CLIP, Vision, or generic prose. An empty SAM 3 object set is a successful no-match result. An invalid Qwen assessment after segmentation stays visible and retryable. A missing board crop is reported as a per-photo failure before Qwen assessment.

This distinction between service-selection fallback and within-operation fallback prevents heterogeneous artifacts and misleading results.

Runtime Lifecycle Summary

stateDiagram-v2
    [*] --> GraphBuilt: construct stable graph
    GraphBuilt --> VisionReady: synchronous initial configuration
    VisionReady --> Checking: refresh installed locations
    Checking --> ProvidersReady: validate bundles and construct providers
    Checking --> PartialAvailability: some bundles missing or invalid
    ProvidersReady --> Configured: publish newer configuration
    PartialAvailability --> Configured: publish explicit capabilities/fallbacks
    Configured --> Running: feature operations
    Running --> Rechecking: download/remove/setting change
    Rechecking --> Configured: identity-diffed apply
    Configured --> [*]: app releases both stable roots

The important invariant is that provider availability may change many times, while application feature identities remain stable for the session.

Source Map

ConcernSource
App retention and startup refreshRawCull/Main/RawCullApp.swift
Model/provider runtimeRawCull/Intelligence/Composition/RawCullAIModelRuntime.swift
Assembly and stable runtimeRawCull/Intelligence/Composition/RawCullIntelligenceRuntime.swift
Runtime paths/capabilitiesRawCull/Intelligence/Contracts/RawCullAIModels.swift
Resource-manager actor/cacheRawCull/Intelligence/ModelManagement/RawCullAIModelResourceManager.swift
Settings refresh/config publicationRawCull/Intelligence/ModelManagement/RawCullAISettingsModel.swift
Downloads and location snapshotRawCull/Intelligence/ModelManagement/RawCullAIModelDownloadsModel.swift
Stable similarity operationsRawCull/Intelligence/Similarity/RawCullSimilarityFeature.swift
Shared similarity/search stateRawCull/Intelligence/Similarity/SimilarityScoringModel.swift
Deep Review service installation/stateRawCull/Intelligence/DeepReview/DeepAIReviewFeature.swift
Qwen provider/model actorRawCull/Intelligence/Qwen/QwenInferenceRuntime.swift
Objects feature, cache reuse, and diagnosticsRawCull/Intelligence/ObjectAnalysis/RawCullObjectAnalysisFeature.swift
Objects contracts, validation, and boardRawCull/Intelligence/ObjectAnalysis
Objects view and statusRawCull/Views/AIAnalysis/ObjectAnalysisView.swift
PhotoAIKit object workflowPhotoAIWorkflows/ObjectSegmentationService.swift

The runtime’s central rule is simple: validate and replace model-dependent services behind stable state owners, then accept results only when revision, generation, catalog, and backend identities still match.