---
title: "Data requirements and standardization"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Data requirements and standardization}
  %\VignetteEncoding{UTF-8}
  %\VignetteEngine{knitr::rmarkdown}
editor_options: 
  markdown: 
    wrap: 80
---

```{r, include = FALSE}
knitr::opts_chunk$set(
  collapse = TRUE,
  comment = "#>",
  eval = TRUE,
  echo = TRUE,
  include = TRUE,
  fig.show = "asis",
  fig.keep = "high",
  warning = FALSE,
  message = FALSE
)
```

## Introduction

The complete workflow is illustrated as follows. This article focuses on the
package's data requirements and the first three steps, including functions
`Mapping()`, `DataCheck()` and `DataStandard()`.

``` text
data -> Mapping() -> DataCheck() -> DataStandard()
     -> prediction / diagnostic / analysis functions
```

```{r}
library(PDRobust)
data("ImperfectConSample", package = "PDRobust")
data("BiSample", package = "PDRobust")
```

Two built-in datasets are included with the package. The datasets are stored in
the `data/` directory. Their generation script is available in `data-raw/` in
the source repository; development scripts are excluded from the CRAN archive.

The first dataset, `BiSample`, is a standardized longitudinal dataset that
satisfies the package’s data requirements. It contains no nonstructural missing
values; outcome values are missing only when they are structurally unobservable
due to any kind of truncation. Each row represents one subject at a specific
time. The dataset includes a subject identifier (`id`), assessment time
(`time`), survival status (`S`), treatment assignment (`A`), a binary outcome
(`Y`), and six subject-level covariates (`X1`–`X6`) .

```{r}
data("BiSample", package = "PDRobust")
head(BiSample)
```

The second one, `ImperfectConSample`, is designed to resemble longitudinal data
collected in a clinical trial with repeated follow-up assessments. Each row
represents a patient observation at a scheduled visit. The dataset includes a
patient identifier (`patient_id`), visit time (`visit_month`), survival status
(`alive_status`), treatment assignment (`treatment`), a continuous clinical
outcome, and six patient-level covariates (`X1`–`X6`) representing relevant
clinical and demographic information.

```{r}
head(ImperfectConSample)
```

## Mapping raw columns and times

`Mapping()` is the sole source of truth for structural columns, analysis window
on the raw time scale, the covariated used to estimate treatment effects, the
variables of interest, and the outcome type.

The arguments `id, time, treatment, survival, outcome` specify the corresponding
column names in the input dataset.

The arguments `baseline_time` and `cutoff_time` define the beginning and end of
the analysis window, respectively. Each must be specified as a single finite
numeric value on the raw time scale, with `baseline_time <= cutoff_time`.
Equal endpoints define a single-time analysis. For example, when observed
time points are (0, 1, ..., 4), the baseline may be set to 0 or 1,
whereas the cutoff may be set to a later time point, such as 4. For
`ImperfectConSample`, the observed time points are (0, 6, 12); therefore, the
analysis window is defined using `baseline_time = 0, cutoff_time = 12`.

The argument `covariates` and `interest_vars` are character vectors containing
the column names of the relevant covariates. Variables specified in
`interest_vars` must also be included in `covariates` . The argument `y_type`
specifies the outcome type. It should be set to `B` for a binary outcome and `C`
for a continuous outcome.

```{r}
map <- Mapping(
  id = "patient_id",
  time = "visit_month",
  treatment = "treatment",
  survival = "alive_status",
  outcome = "clinical_outcome",
  baseline_time = 0,
  cutoff_time = 12,
  covariates = paste0("X", 1:6),
  interest_vars = c("X1", "X2"),
  y_type = "C" # "B"
)

```

Users can inspect the mapping details using `print()` and the `attributes()`
function may additionally be used to inspect object-level metadata, such as its
class.

```{r}
print(map)
```

```{r}
attributes(map)
```

## Validation without modification

`DataCheck()` evaluates whether the input dataset satisfies the structural and
analytical requirements of `PDRobust` without altering the data. It identifies
potential issues, reports their severity and recommended handling, and
determines whether the dataset is ready for analysis, can be standardized using
`DataStandard()`, or requires manual resolution.

Each check includes an action-oriented message. When `strict = FALSE`, which is
the default, `DataCheck()` returns a validation report without
modifying the input dataset. When `strict = TRUE`, the function raises an error
if one or more failed checks are marked as analysis-blocking.
Missing required columns or empty data cause an early return containing only
the checks that can be performed at that stage.

```{r}
check <- DataCheck(ImperfectConSample, map, strict = FALSE)

attributes(check)
```

`DataCheck()` returns an object of class `pd_data_check`. The object contains
the following components:

| Component | Description |
|----|----|
| `valid` | Indicates whether all checks with severity `"error"` have passed. Informational messages and warnings do not by themselves make the report invalid. |
| `ready_for_analysis` | `TRUE` when no failed check is marked as analysis-blocking. Raw data still need `DataStandard()` to attach the mapping and readiness metadata required by the HTE and diagnostic interfaces. |
| `manual_resolution_required` | Indicates whether at least one failed check requires manual review or correction. Such issues are not automatically resolved by `DataStandard()`. |
| `can_standardize` | `TRUE` when no failed check requires manual resolution. Deletion may still require `drop = TRUE`, leave no observations, or remove a treatment group, so this flag does not guarantee success or final analysis readiness. |
| `checks` | A data frame containing the itemized validation results, including the status, severity, diagnostic summary, analysis implications, and recommended handling for each check. |
| `settings` | Records the settings used during validation. The current implementation stores the validated mapping object in `settings$mapping`. |
| `diagnostics` | Contains detailed supporting information, such as affected row numbers, subject identifiers, missingness summaries, treatment-group counts, and problematic covariates. |

```{r}
check$valid
check$ready_for_analysis
check$manual_resolution_required
check$can_standardize
```

The following are diagnostics of check for `ImperfectConSample`. Users can find
detailed descriptions of all validation items in the article
*Details-for-DataCheck*.

```{r}
head(check$diagnostics)
```

## Standardization and audit attributes

`DataStandard()` returns a prepared data for later analysis.

```{r}
pd_data <- DataStandard(ImperfectConSample, map, drop = TRUE)
class(pd_data)
head(pd_data)
```

The returned `pd_data` object contains several attributes that document its
structure and the transformations applied during standardization. These
attributes can be inspected using:

```{r}
names(attributes(pd_data))
```

| Attribute | Description |
|:---|:---|
| `names` | Stores the column names of the standardized data frame. |
| `row.names` | Stores the row identifiers used by the data frame. |
| `class` | Identifies the object classes, including its data-frame and package-specific classes. |
| `pd_mapping` | Stores the mapping used by subsequent package functions. Column names retain their input names, while baseline and cutoff times refer to the standardized grid. |
| `pd_original_mapping` | Preserves the original user-supplied mapping on the raw data scale before identifiers, time points, and encodings were standardized. |
| `pd_check` | Stores the validation results associated with the standardized dataset, including readiness indicators, detected issues, and recommended handling. |
| `pd_standardization` | Contains `time_map`, `id_map`, `attrition`, and `initial_check`, recording identifier/time conversions, exclusion counts and subject-level reasons, and the input validation report. |

For exmaple, the original and standardized mappings can be compared using:

```{r}
attr(pd_data, "pd_original_mapping")
attr(pd_data, "pd_mapping")
```

Among these attributes, `pd_standardization` is particularly important because
it provides the primary audit trail for the changes made to the input dataset.
It can be inspected directly using:

```{r}
standardization <- attr(pd_data, "pd_standardization")
names(standardization)
```

It retains all observed assessment times within the analysis window defined by
the mapping object and transforms the ordered time grid to consecutive integers
(0, 1, ..., n).

```{r}
standardization$time_map
```

Subject identifiers are similarly mapped to consecutive integers, explicitly
recognized binary encodings are converted safely, and the resulting longitudinal
dataset is sorted by subject and standardized analysis time.

For a subject to be retained, the dataset must contain one usable record at each
retained assessment time.

```{r}
head(standardization$id_map)
```

Rows outside the mapped time window are always removed. With `drop = TRUE`,
rows with missing identifiers or times can also be removed; remaining subjects
with missing visits or required analysis values are excluded in full. The
attrition report records the row-removal counts and subject-level exclusions.
Outcome values that are structurally unobservable after death or other trunction
are distinguished from ordinary missing outcomes among surviving subjects and
are therefore handled separately during validation and standardization.

```{r}
standardization$attrition
```

Standardization performs a second validation on the retained data. If dropping
subjects removes a treatment group, it returns a warning and a `pd_check`
attribute with `ready_for_analysis = FALSE`; the HTE interfaces reject that
object. Inspect this flag before continuing:

```{r}
attr(pd_data, "pd_check")$ready_for_analysis
```

The ID audit map contains one row per retained subject, and detailed reports
can contain row or subject indices. These attributes therefore grow with the
data and the number of detected problems. They describe the standardization
call; subsequent editing or subsetting does not recompute them. Revalidate
changed data before analysis.

Finally, even when a built-in dataset such as `BiSample`, or a user-supplied
dataset, already satisfies all `PDRobust` data requirements, it should still be
processed through the package’s data-preparation workflow before analysis. This
ensures that the dataset is formally validated, standardized, and supplied with
the mapping and audit attributes required by downstream functions.
