What "Large Shoot" Means

The edit prep workflow scales differently with footage volume than most editors expect. The challenges that emerge at 100+ clips are different from the challenges at 30 clips, and the workflow needs to adapt accordingly.

For this guide, a large shoot means roughly:

  • 100 or more individual media files
  • Multiple cameras (often 3-5+ angles)
  • External audio recorders alongside camera audio
  • 10-30+ hours of total footage
  • Mixed content types: dialogue, B-roll, action, environmental

This describes typical multi-day branded shoots, mid-budget documentary shoots, multi-camera live events, and conference recordings. It is not feature documentary territory (where you might have hundreds of hours over months) but it is well past the simple single-camera shoot territory.

The workflow below assumes a Premiere Pro target NLE, but the principles apply to FCP and DaVinci Resolve. The AI tooling integrates similarly across these NLEs because each supports project file import and bin organization.

EDITOR'S TAKE

The biggest mistake I see on large shoots is editors trying to skip prep stages because the project feels urgent. "Let's just start cutting" works on a 30-clip project. On a 200-clip project, skipping prep means spending the entire edit lost in the bins. The discipline to do the prep correctly upfront is the single biggest predictor of whether a large shoot edits smoothly or becomes a disaster.

Stage 1: Ingest and Backup

Ingest is the mechanical work of getting media off cards and onto editing storage. It sounds simple but is often where avoidable problems start, because lost or duplicated files at this stage cause cascading issues in later stages.

Use a checksum-verified copy tool. Tools like Hedge, Shotput Pro, or Silverstack copy files with byte-level verification, ensuring that what arrives on your storage is bit-identical to what was on the card. This catches corruption that simple drag-and-drop copies might miss.

Copy to two destinations simultaneously. The standard practice is to copy each card to both your editing storage and an archive drive in parallel. This produces an immediate backup before you start editing.

Preserve original folder structures from cards. Camera cards have specific folder structures (DCIM, BPAV, AVCHD) that NLEs sometimes need to import the media correctly. Do not flatten or rename inside the card structure during ingest -- preserve it and put it inside your project's organization layer.

Verify before deleting cards. Confirm checksums match and spot-check files in your NLE before reformatting any cards. The discipline of "never erase the source until verified" prevents catastrophic data loss.

Time for a 100-clip ingest with 10-20 hours of footage: 1-3 hours depending on transfer speeds. Most of this is wall-clock time during which you can do other work. The active editor time is 15-30 minutes of setup and verification.

Where AI helps: Not much at this stage. Ingest is mechanical and AI does not meaningfully accelerate it. The next stages benefit from AI more.

Stage 2: Folder Organization

Once footage is ingested, organize it into a structure that will support the rest of the workflow. This organization persists across the entire project, so getting it right pays back across every subsequent stage.

A workable structure for large shoots:

  • 01_Source/ -- Untouched original media organized by source
    • Camera_A/
    • Camera_B/
    • Camera_C/
    • Audio_External/
    • Drone/
  • 02_Proxies/ -- Generated proxy files (if applicable)
  • 03_Audio/ -- Music, sound effects, voiceover
  • 04_Graphics/ -- Logos, lower thirds templates, motion graphics
  • 05_Project/ -- NLE project files
  • 06_Renders/ -- Output renders for review and delivery
  • 07_Documents/ -- Scripts, briefs, transcripts, notes

This structure is hierarchical and self-explanatory. Anyone joining the project can navigate it without instructions. The numbered prefix forces the folders into a logical order that matches the workflow.

Within each camera folder, organize by date and scene if shooting spanned multiple days or scenes. "Day1_Sc1_FounderInterview" is more useful than "Cards_1_through_4." The camera operators usually know what was shot when, so use that knowledge while it is fresh.

Where AI helps: Some AI tools can analyze ingested footage and propose organizational tags automatically ("this card contains interior interview footage," "this card contains exterior B-roll"). The proposals are not always right but can speed initial organization. Editor verification is needed.

Time: 30-90 minutes for a 100-clip shoot.

Stage 3: Multi-Source Sync

Multi-camera shoots need synced audio across cameras, plus alignment with external audio recorders if used. This is one of the most time-consuming traditional prep tasks and one of the biggest AI wins.

Traditional sync workflow: identify a sync reference (clap, slate, sync beep), find the same reference in each source's audio waveform, manually align them, repeat for every multi-camera scene. For a shoot with 8 multi-camera scenes, this is typically 60-90 minutes of focused work.

AI sync workflow: feed all sources to the AI tool. It analyzes audio waveforms across sources and identifies matching audio patterns automatically. Sync happens in seconds rather than minutes per scene.

Modern AI sync handles common cases reliably:

  • Cameras with matching scratch audio
  • Cameras synced to external timecode
  • External audio recorders alongside cameras
  • Mixed sample rates between sources

Edge cases that still need manual attention:

  • Scenes where camera audio is heavily processed (noise reduction, EQ)
  • Cameras that lack scratch audio entirely
  • Long takes where drift between unsynced cameras compounds
  • Mixed frame rates across cameras

For an 8-scene multi-camera shoot, AI sync takes 5-15 minutes total instead of 60-90 minutes manually. The 4-5x time savings here is one of the easiest AI wins on large shoots.

Where AI helps: Significantly. Audio waveform matching is a well-defined ML problem and modern tools solve it reliably.

Stage 4: Transcription and Indexing

For dialogue-heavy shoots (interviews, panels, presentations), transcription is foundational to everything that follows. Once transcripts exist, you can search footage by content, find specific moments instantly, and build selects from transcript snippets rather than from scrubbing.

Traditional transcription is brutal. A 60-minute interview takes 3-6 hours to transcribe by hand at editing-quality accuracy. Outsourcing to a transcription service is faster (24-48 hour turnaround) but adds cost and waiting time.

AI transcription happens in minutes at production quality:

  • Clean studio audio: 95-97% accuracy
  • Outdoor/noisy audio: 85-92% accuracy
  • Multiple speakers: speaker labels generally correct
  • Multiple languages: increasingly supported with good accuracy

For a 10-hour interview shoot, AI transcription completes in 30-90 minutes of background processing. Manual transcription would take 30-60 hours. The order-of-magnitude difference is what makes the rest of the workflow possible.

Beyond transcription, modern AI tools index the footage along multiple dimensions:

  • Speech transcript with timecodes and speaker labels
  • Visual content tags (shot type, subject, action)
  • Audio metadata (silence, music, ambient noise)
  • Speaker identification across multicam

This combined index is searchable. "Find all clips where the founder mentions pricing" returns ranked results in seconds. "Find wide shots with movement" returns visual matches. "Find moments with laughter" returns audio matches.

Where AI helps: Massively. This stage is the single biggest time saver in the entire workflow.

Time: 30-90 minutes of background processing plus 15-30 minutes of editor verification.

Stage 5: Logging and Tagging

Logging is the work of describing what is in each clip with enough precision to find it later. On large shoots this is genuinely time-consuming traditional work and very high-value AI work.

Traditional logging: an assistant editor watches every clip, writes a brief description, marks usable takes, notes technical issues. For 100 clips averaging 5-10 minutes each, this is 8-15 hours of focused viewing and typing.

AI logging: the AI tool generates clip-level metadata automatically based on its visual and audio analysis. Each clip gets:

  • Shot type (wide, medium, close, OTS, etc.)
  • Subject (who or what is in the frame)
  • Action (what is happening)
  • Audio characteristics (dialogue, music, ambient)
  • Quality flags (focus issues, audio issues, motion issues)
  • Speaker if present (for dialogue)

The editor reviews and refines the AI tags rather than generating them from scratch. Most AI tags are correct (85-95%) and the work shifts to verifying and adding clip-specific notes the AI cannot infer ("this is the take where the founder gets emotional," "this take has a generator hum starting at 2:30").

For 100 clips, AI-assisted logging takes 1-2 hours instead of 8-15 hours manually. The savings come from the AI doing 80-90 percent of the descriptive work and the editor handling the high-judgment 10-20 percent.

AI-ASSISTED LOGGING WORKFLOW
01
Run Bulk Auto-Tagging
Let the AI tool generate descriptive tags for every clip in the project. This runs in the background.
02
Spot-Check Tag Accuracy
Review tags on 10-20 random clips to verify quality. If accuracy is high, trust most tags. If low, investigate why (codec issues, unusual content type) before proceeding.
03
Add High-Judgment Notes
Add notes the AI couldn't infer: emotional moments, technical issues specific to this shoot, references to the script or brief, take quality opinions.
04
Mark Selects
Identify candidate selects (best takes, strongest moments). This is the bridge into the next stage and benefits from human judgment.

Where AI helps: Significantly on the descriptive layer; less on the judgment layer (which is where it should help less).

Stage 6: Building Selects

Selects are clips identified as strong candidates for use in the final edit. They are the working set the editor builds the rough cut from. On large shoots, selects discipline is critical because trying to navigate 100+ raw clips during the edit is overwhelming.

Traditional selects workflow: editor watches everything, marks promising clips, copies marked clips into selects bins organized by topic or scene. This is judgment-intensive work that takes 4-10 hours on a large shoot.

AI-assisted selects workflow: the editor uses search to find candidates by content, then makes select decisions on each candidate. The search makes the discovery phase fast; the editor's judgment makes the selection decisions sound.

Practical workflow:

Search by intent. "Show me the strongest moments where the founder talks about origin." "Find all clips with the customer smiling." "Find wide establishing shots of the office." Each search returns ranked candidates.

Review candidates quickly. The AI's ranking gets the strongest candidates near the top. The editor reviews 5-15 candidates per search and marks 1-3 as selects.

Build selects bins thematically. Organize selects by topic, scene, or function rather than by camera or chronology. "Founder Origin Selects," "Product Demo Selects," "Customer Reactions Selects" -- whatever maps to your intended structure.

Avoid over-selecting. The temptation is to mark too many clips as selects, which defeats the purpose. Aim for selects to be 3-5x your finished length, not 10x. Tighter selects produce faster rough cuts.

For a 10-hour shoot targeting a 5-minute finished video, expect to identify 15-25 minutes of selects across maybe 30-50 clips. AI-assisted, this takes 1-2 hours instead of 4-10 hours manually.

Where AI helps: Discovery is dramatically faster; selection decisions remain editor-driven.

Stage 7: Stringouts

Stringouts are the final prep deliverable: long sequences containing every selected clip in approximate order, organized so the editor can begin building the rough cut from a structured working set.

For dialogue-heavy interview content, stringouts are typically organized by topic or speaker. All the founder origin content, then all the founder vision content, then all the customer testimonial content, etc. Within each topic, clips are arranged in approximate order of intended use, with stronger candidates earlier.

For action or B-roll-heavy content, stringouts are organized by scene or location. All the office interior shots, all the exterior establishing shots, all the product close-ups, etc.

The stringout is not the rough cut. It is too long, too uncurated, and lacks structural commitment. But it is dramatically faster to navigate than raw bins, which makes the rough cut workflow itself much faster.

AI assists with stringout building by automatically organizing selects into thematic groups based on transcript content and visual tags. The editor reviews the AI's grouping, refines it, and within each group orders clips by judgment (not by AI ranking, which is not yet reliable on "which clip should come first").

For a typical large-shoot project, stringouts take 30-90 minutes with AI assistance vs 2-4 hours manually. The savings come from automated grouping; the editor still spends judgment time on intra-group ordering.

Where AI helps: Moderately. Organization is automatable; ordering is partially automatable.

Total Time and Where AI Helps

Stacking up all seven stages, the total prep workflow time on a typical 100-clip, 10-hour multi-camera shoot looks like this:

StageManual TimeAI-Assisted TimeReduction
1. Ingest and backup1-3 hr1-3 hrMinimal
2. Folder organization30-90 min20-60 min30%
3. Multi-source sync60-90 min5-15 min85%
4. Transcription / indexing15-30 hr30-90 min95%
5. Logging and tagging8-15 hr1-2 hr85%
6. Building selects4-10 hr1-2 hr75%
7. Stringouts2-4 hr30-90 min65%
Total30-55 hr4-9 hr~85%

The aggregate reduction is dramatic because AI compresses multiple stages, and the savings compound. The biggest individual wins are in transcription and logging, which are the most repetitive and AI-friendly tasks. Sync is also a large relative win because audio matching is a well-defined ML problem.

The remaining 4-9 hours of editor time go into the parts that genuinely require human judgment: reviewing AI outputs, adding context the AI cannot infer, building selects with editorial taste, and ordering stringouts based on intent. These are not the parts that benefit from compression -- they are the parts where editor time creates the most value.

The end state of this workflow is a project ready to begin rough cut work. The editor knows what footage exists, where it lives, what is in it, and which clips are worth using. From there, the rough cut workflow takes over -- which itself benefits from AI in similar ways. For the next stage, see our walkthrough of creating a rough cut with AI and our breakdown of logging and labeling footage with AI.

TRY IT

Stop scrubbing. Start creating.

Wideframe gives your team an AI agent that searches, organizes, and assembles Premiere Pro sequences from your footage. 7-day free trial.

REQUIRES APPLE SILICON

Frequently asked questions

Edit prep on large shoots breaks into seven sequential stages: ingest and backup, folder organization, multi-source sync, transcription and indexing, logging and tagging, building selects, and stringouts. Each stage produces a deliverable that supports the next. AI compresses every stage by 50-95% by automating the mechanical layers.

Manually, edit prep on a 100-clip multi-camera shoot with 10 hours of footage takes 30-55 hours. With AI assistance -- automated transcription, sync, logging, and search -- the same workflow completes in 4-9 hours. The biggest savings come from transcription, logging, and multi-source sync.

A numbered hierarchical structure works well: 01_Source organized by camera and external audio, 02_Proxies, 03_Audio, 04_Graphics, 05_Project, 06_Renders, 07_Documents. The numbered prefix forces logical order. Within source folders, organize by date and scene rather than by card number.

AI for routine multi-camera sync where cameras share scratch audio. Manual for edge cases -- heavily processed audio, cameras lacking scratch audio, or long takes with drift. AI sync handles 90%+ of typical large-shoot scenes correctly and reduces sync time by about 85%.

Selects are individual clips identified as strong candidates for use in the final edit. Stringouts are long sequences containing all those selects arranged by topic or scene in approximate order. Selects are the curated set; stringouts are the structured working timeline that feeds into rough cut work.

DP
Daniel Pearson
Co-Founder & CEO, Wideframe
Daniel Pearson is the co-founder & CEO of Wideframe. Before founding Wideframe, he founded an agency that made thousands of video ads. He has a deep interest in the intersection of video creativity and AI. We are building Wideframe to arm humans with AI tools that save them time and expand what's creatively possible for them.
This article was written with AI assistance and reviewed by the author.