DataModule#

class pyqit.DataModule(X, y, name: str = 'dataset', normalize: str | None = None, split: tuple = (0.7, 0.15, 0.15), stratify: bool = False, seed: int | None = 42, batch_size: int = 32, num_workers: int = 0, transform: Callable | list | None = None, shuffle: bool = True, drop_last: bool = False)[source]#

Bases: object

Lazy split, normalize, and prescale for a classical dataset.

Nothing runs until setup(), which Trainer.fit/.predict call for you. Order is split, then stateful normalization fit on train only, then stateless quantum prescaling driven by the model’s embedding.

Parameters:
  • X (array-like)

  • y (array-like)

  • name (str, default "dataset")

  • normalize ({"minmax", "zscore", "l2", "l1"}, optional)

  • split (tuple of float, default (0.70, 0.15, 0.15)) – Train, val, test fractions. Must sum to 1.0.

  • stratify (bool, default False)

  • seed (int, optional, default 42)

  • batch_size (int, default 32)

  • num_workers (int, default 0) – Torch backend only.

  • transform (callable or list of callable, optional)

  • shuffle (bool, default True)

  • drop_last (bool, default False)

Examples

>>> import pyqit
>>> dm = pyqit.DataModule(X, y, normalize="minmax", batch_size=16)
>>> history = pyqit.Trainer(max_epochs=10).fit(model, dm)
property X_test#

Test split features, or None with no test split. Raises before setup.

property X_train#

Train split features. Raises before setup().

property X_val#

Val split features, or None with no val split. Raises before setup().

property class_labels#

Unique values in y.

clone_empty() → DataModule[source]#

Return an unsetup shallow copy holding a copy of the fitted normalizer.

property feature_dim#

Feature count after setup(); falls back to n_features before it.

for_prediction(X) → DataModule[source]#

A predict-only DataModule over new raw rows, sharing this one’s fitted state.

Keeps the fitted normalizer, encoder and n_qubits, so Trainer.predict processes X exactly as it would this DataModule’s own test split.

Parameters:

X (array-like) – Raw rows, the same kind of input this DataModule was built from.

Returns:

Not yet set up; Trainer.predict sets it up with stage=”predict”.

Return type:

DataModule

Examples

>>> preds = trainer.predict(model, dm.for_prediction(X_new))
classmethod from_csv(path: str | Path, label_col: str | int = -1, delimiter: str = ',', **kw) → DataModule[source]#

Build from a CSV file. See from_dataframe for label_col.

classmethod from_dataframe(df: Any, label_col: str | int = -1, **kw) → DataModule[source]#

Build from a DataFrame, splitting off label_col as y.

Parameters:
  • df (pandas.DataFrame)

  • label_col (str or int, default -1) – Column name, or a position (negative indexes from the end).

  • **kw – Forwarded to the constructor.

classmethod from_numpy(X, y, **kw)[source]#

Build like the constructor; kept for a consistent from_* family.

classmethod from_sklearn(loader: Callable, **kw) → DataModule[source]#

Build from an sklearn dataset loader.

Parameters:
  • loader (callable) – E.g. sklearn.datasets.load_iris.

  • **kw – Forwarded to the constructor.

Examples

>>> from sklearn.datasets import load_iris
>>> dm = pyqit.DataModule.from_sklearn(load_iris)
classmethod load(path: str, X, y=None) → DataModule[source]#

A DataModule over X with the settings and preprocessing in path.

Parameters:
  • path (str) – A file written by save.

  • X (array-like) – Raw rows, the same kind of input the saved DataModule was built from.

  • y (array-like, optional) – Targets. Without them the result is predict-only over X, as for_prediction returns, and Trainer.predict applies the saved normalizer statistics rather than refitting. With them the saved split is kept and setup refits the normalizer on the new rows.

Returns:

Not yet set up.

Return type:

DataModule

property n_classes#

Number of unique values in y.

property n_features#

Raw feature count, before any prescaling.

property n_samples#

Total rows in X, before splitting.

property normalizer: _Normalizer | None#

The fitted _Normalizer, or None before setup or with no normalize.

reconfigure(**kwargs) → DataModule[source]#

Update settings and clear fitted state, requiring a re-setup().

Parameters:

**kwargs – Any of normalize, split, stratify, seed, batch_size, num_workers, transform.

Returns:

self.

Return type:

DataModule

save(path: str) → str[source]#

Pickle the settings and fitted preprocessing, without the data.

Writes what for_prediction carries over: every constructor setting, transform included, the fitted normalizer, encoder and n_qubits. The raw arrays and splits are dropped, so the file is small and load needs new rows. Loading a pickle runs code from the file, so load only files you wrote, the same rule as torch.load.

Parameters:

path (str) – File to write. Parent directories are created.

Returns:

path.

Return type:

str

setup(stage: str | None = None, batch_size: int | None = None, n_qubits: int | None = None, encoder: type | None = None, force: bool = False) → DataModule[source]#

Split, normalize, and prescale. Idempotent unless force=True.

Trainer.fit/.predict call this for you, passing n_qubits and encoder from the model so quantum prescaling is never skipped.

Parameters:
  • stage ({"fit", "val", "test", "predict"}, optional) – “predict” skips the split and uses the whole dataset as test. Every other stage splits.

  • batch_size (int, optional) – Applied even when already set up.

  • n_qubits (int, optional) – From the model. Persists across a later force=True call that omits it.

  • encoder (type, optional) – Embedding class. Same persistence as n_qubits.

  • force (bool, default False) – Redo the split even if already set up.

Returns:

self.

Return type:

DataModule

property splits#

(X_train, y_train, X_val, y_val, X_test, y_test).

test_loader(shuffle: bool = False)[source]#

Build a loader like train_loader, over test. None with no test split.

to_lightning()[source]#

Converts itself to a Lightning adapter if the backend requires it.

train_loader(shuffle: bool | None = None, drop_last: bool | None = None)[source]#

Build a DataLoader (torch) or _NumpyLoader (pennylane) over train.

shuffle and drop_last default to this DataModule’s own settings.

val_loader(shuffle: bool = False)[source]#

Build a loader like train_loader, over val. None with no val split.

property y_test#

Test split targets, or None with no test split. Raises before setup().

property y_train#

Train split targets. Raises before setup().

property y_val#

Val split targets, or None with no val split. Raises before setup().