DataModule#
- class pyqit.DataModule(X, y, name: str = 'dataset', normalize: str | None = None, split: tuple = (0.7, 0.15, 0.15), stratify: bool = False, seed: int | None = 42, batch_size: int = 32, num_workers: int = 0, transform: Callable | list | None = None, shuffle: bool = True, drop_last: bool = False)[source]#
Bases:
objectLazy split, normalize, and prescale for a classical dataset.
Nothing runs until setup(), which Trainer.fit/.predict call for you. Order is split, then stateful normalization fit on train only, then stateless quantum prescaling driven by the model’s embedding.
- Parameters:
X (array-like)
y (array-like)
name (str, default "dataset")
normalize ({"minmax", "zscore", "l2", "l1"}, optional)
split (tuple of float, default (0.70, 0.15, 0.15)) – Train, val, test fractions. Must sum to 1.0.
stratify (bool, default False)
seed (int, optional, default 42)
batch_size (int, default 32)
num_workers (int, default 0) – Torch backend only.
transform (callable or list of callable, optional)
shuffle (bool, default True)
drop_last (bool, default False)
Examples
>>> import pyqit >>> dm = pyqit.DataModule(X, y, normalize="minmax", batch_size=16) >>> history = pyqit.Trainer(max_epochs=10).fit(model, dm)
- property X_test#
Test split features, or None with no test split. Raises before setup.
- property X_train#
Train split features. Raises before setup().
- property X_val#
Val split features, or None with no val split. Raises before setup().
- property class_labels#
Unique values in y.
- clone_empty() DataModule[source]#
Return an unsetup shallow copy holding a copy of the fitted normalizer.
- property feature_dim#
Feature count after setup(); falls back to n_features before it.
- for_prediction(X) DataModule[source]#
A predict-only DataModule over new raw rows, sharing this one’s fitted state.
Keeps the fitted normalizer, encoder and n_qubits, so Trainer.predict processes X exactly as it would this DataModule’s own test split.
- Parameters:
X (array-like) – Raw rows, the same kind of input this DataModule was built from.
- Returns:
Not yet set up; Trainer.predict sets it up with stage=”predict”.
- Return type:
Examples
>>> preds = trainer.predict(model, dm.for_prediction(X_new))
- classmethod from_csv(path: str | Path, label_col: str | int = -1, delimiter: str = ',', **kw) DataModule[source]#
Build from a CSV file. See from_dataframe for label_col.
- classmethod from_dataframe(df: Any, label_col: str | int = -1, **kw) DataModule[source]#
Build from a DataFrame, splitting off label_col as y.
- Parameters:
df (pandas.DataFrame)
label_col (str or int, default -1) – Column name, or a position (negative indexes from the end).
**kw – Forwarded to the constructor.
- classmethod from_numpy(X, y, **kw)[source]#
Build like the constructor; kept for a consistent from_* family.
- classmethod from_sklearn(loader: Callable, **kw) DataModule[source]#
Build from an sklearn dataset loader.
- Parameters:
loader (callable) – E.g. sklearn.datasets.load_iris.
**kw – Forwarded to the constructor.
Examples
>>> from sklearn.datasets import load_iris >>> dm = pyqit.DataModule.from_sklearn(load_iris)
- classmethod load(path: str, X, y=None) DataModule[source]#
A DataModule over
Xwith the settings and preprocessing inpath.- Parameters:
path (str) – A file written by save.
X (array-like) – Raw rows, the same kind of input the saved DataModule was built from.
y (array-like, optional) – Targets. Without them the result is predict-only over
X, as for_prediction returns, and Trainer.predict applies the saved normalizer statistics rather than refitting. With them the saved split is kept andsetuprefits the normalizer on the new rows.
- Returns:
Not yet set up.
- Return type:
- property n_classes#
Number of unique values in y.
- property n_features#
Raw feature count, before any prescaling.
- property n_samples#
Total rows in X, before splitting.
- property normalizer: _Normalizer | None#
The fitted _Normalizer, or None before setup or with no normalize.
- reconfigure(**kwargs) DataModule[source]#
Update settings and clear fitted state, requiring a re-setup().
- Parameters:
**kwargs – Any of normalize, split, stratify, seed, batch_size, num_workers, transform.
- Returns:
self.
- Return type:
- save(path: str) str[source]#
Pickle the settings and fitted preprocessing, without the data.
Writes what for_prediction carries over: every constructor setting, transform included, the fitted normalizer, encoder and n_qubits. The raw arrays and splits are dropped, so the file is small and load needs new rows. Loading a pickle runs code from the file, so load only files you wrote, the same rule as
torch.load.- Parameters:
path (str) – File to write. Parent directories are created.
- Returns:
path.- Return type:
str
- setup(stage: str | None = None, batch_size: int | None = None, n_qubits: int | None = None, encoder: type | None = None, force: bool = False) DataModule[source]#
Split, normalize, and prescale. Idempotent unless force=True.
Trainer.fit/.predict call this for you, passing n_qubits and encoder from the model so quantum prescaling is never skipped.
- Parameters:
stage ({"fit", "val", "test", "predict"}, optional) – “predict” skips the split and uses the whole dataset as test. Every other stage splits.
batch_size (int, optional) – Applied even when already set up.
n_qubits (int, optional) – From the model. Persists across a later force=True call that omits it.
encoder (type, optional) – Embedding class. Same persistence as n_qubits.
force (bool, default False) – Redo the split even if already set up.
- Returns:
self.
- Return type:
- property splits#
(X_train, y_train, X_val, y_val, X_test, y_test).
- test_loader(shuffle: bool = False)[source]#
Build a loader like train_loader, over test. None with no test split.
- train_loader(shuffle: bool | None = None, drop_last: bool | None = None)[source]#
Build a DataLoader (torch) or _NumpyLoader (pennylane) over train.
shuffleanddrop_lastdefault to this DataModule’s own settings.
- val_loader(shuffle: bool = False)[source]#
Build a loader like train_loader, over val. None with no val split.
- property y_test#
Test split targets, or None with no test split. Raises before setup().
- property y_train#
Train split targets. Raises before setup().
- property y_val#
Val split targets, or None with no val split. Raises before setup().