Batch processing

library(tidymedia)

tidymedia is made for running the same job on many files. The examples on this page use run = FALSE. Each function then returns its FFmpeg commands without running them, so you can read them first. Leave out run = FALSE to process the files.

folder <- system.file("extdata", package = "tidymedia")
video <- system.file("extdata", "sample.mp4", package = "tidymedia")

The batch runner

ffm_batch() runs a job for each row of a jobs table. You also give it a function, .f, that turns one row into a pipeline. Each column of the table goes to .f as an argument of the same name, as in purrr::pmap().

ffm_batch() returns your jobs table with a command column added. When run = TRUE, it also adds a success column.

ffm_jobs() makes a jobs table from a folder. It has one row for each file of the media type you ask for, with the file’s full path in an input column. If the folder has no such files, ffm_jobs() stops with an error.

Add any other columns that .f needs. Here each input gets an output:

jobs <- ffm_jobs(folder, type = "video")
jobs$output <- paste0(tools::file_path_sans_ext(basename(jobs$input)), ".mp3")

ffm_batch(jobs, run = FALSE, .f = function(input, output, ...) {
  ffm_files(input, output) |>
    ffm_drop("video") |>
    ffm_codec(audio = "libmp3lame")
})
#> # A tibble: 1 × 3
#>   input                                                           output command
#>   <chr>                                                           <chr>  <chr>  
#> 1 /private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI… sampl… "-y -i…

You write the pipeline, so each job can use any pipeline functions. Give .f a ... argument, so that it accepts table columns it does not use. Without ..., such a column stops the batch with R’s “unused argument” error.

Batch task functions

For common jobs, you do not need to write .f. Each task function has a *_batch() version that takes a jobs table and runs the task on each row. Examples are extract_audio_batch(), convert_audio_batch(), crop_video_batch(), standardize_video_batch() and normalize_audio_batch().

Some batch functions, such as crop_video_batch(), can take the table from ffm_jobs() as it is. Others, such as extract_audio_batch(), need an output column first:

jobs <- ffm_jobs(folder, type = "video")
jobs$output <- paste0(tools::file_path_sans_ext(basename(jobs$input)), "_cropped.mp4")

crop_video_batch(jobs, width = 160, height = 120, run = FALSE)
#> # A tibble: 1 × 3
#>   input                                                           output command
#>   <chr>                                                           <chr>  <chr>  
#> 1 /private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI… sampl… "-y -i…

Without an output column, crop_video_batch() adds _cropped to each input name and writes to the input’s folder. Here that folder is inside the installed package. So the example adds an output column that writes to the working folder instead.

crop_video_batch() stops with an error if two rows would write the same output file. vignette("workflow") uses several batch functions on a study folder.

One input, many outputs

Some tasks make many outputs from one input.

segment_video() cuts a file into pieces at the start and end times you give. It returns one row for each piece:

segment_video(
  video,
  start = c(0, 0.5),
  end  = c(0.5, 1),
  run = FALSE
)
#> # A tibble: 2 × 5
#>   input                                               output start   end command
#>   <chr>                                               <chr>  <dbl> <dbl> <chr>  
#> 1 /private/var/folders/px/frfvbz4n0sx90__c62fwwzz400… /priv…   0     0.5 "-y -i…
#> 2 /private/var/folders/px/frfvbz4n0sx90__c62fwwzz400… /priv…   0.5   1   "-y -i…

separate_audio_video() writes the audio and the video of a file to two files. It returns the two commands:

separate_audio_video(video, "audio.aac", "video.mp4", run = FALSE)
#>                                                                                                                                                                   audio 
#> "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -codec:a copy -map \"0:a\" \"audio.aac\"" 
#>                                                                                                                                                                   video 
#> "-y -i \"/private/var/folders/px/frfvbz4n0sx90__c62fwwzz40000gn/T/RtmpI5eZw9/Rinst877f72a946bf/tidymedia/extdata/sample.mp4\" -codec:v copy -map \"0:v\" \"video.mp4\""

Running in parallel

These functions take parallel = TRUE:

separate_audio_video() does not take it, but separate_audio_video_batch() does. On probe_container(), probe_streams(), probe_video() and probe_audio(), the argument has an effect only when you pass infile.

With parallel = TRUE, the jobs run through furrr. They run in parallel only if you set a future plan. With no plan, the jobs run one at a time, and R gives a warning that says so:

library(future)
plan(multisession)

ffm_batch(jobs, parallel = TRUE, .f = function(input, output, ...) {
  ffm_files(input, output) |> ffm_drop("video") |> ffm_codec(audio = "libmp3lame")
})

The result has the command for each job. Save that column, and you have a full record of the FFmpeg commands that made your files.

Where to next