Zum Inhalt

sources

module wittgenstein_catalog_builder.sources

Source-extraction strategies (Adapter pattern) feeding CatalogBuilder.

CatalogBuilder infers a catalog from a DataFrame that already exists in the caller's process — until now, the caller was on the hook for loading that DataFrame itself. SourceStrategy is the target interface (the same Adapter shape wittgenstein-crm-client already establishes: one ABC, concrete strategies implementing it, no registry/factory); a caller imports and instantiates the concrete strategy it wants directly, and CatalogBuilder.build_from_strategy extracts + infers in one call.

Classes

wittgenstein_catalog_builder.sources.SourceStrategy

class SourceStrategy()

Bases : ABC

A named tabular source CatalogBuilder can extract + infer from.

Methods

  • extract — Load the full source as one DataFrame.

  • extract_chunks — Load the source in row-chunks of at most chunk_size.

wittgenstein_catalog_builder.sources.SourceStrategy.extract

method SourceStrategy.extract() → pd.DataFrame

Load the full source as one DataFrame.

wittgenstein_catalog_builder.sources.SourceStrategy.extract_chunks

method SourceStrategy.extract_chunks(chunk_size: int) → Iterator[pd.DataFrame]

Load the source in row-chunks of at most chunk_size.

The default splits the result of extract() in memory — fine for sources small enough to already fit in memory (Excel/CSV). A source with its own server-side pagination (e.g. a large SAP view) should override this instead of loading everything up front.

wittgenstein_catalog_builder.sources.ExcelCsvSourceStrategy

class ExcelCsvSourceStrategy(path: Path, *, sheet_name: str | None = None)

Bases : SourceStrategy

Wraps the existing Excel/CSV read+clean helpers — zero behavior change.

An existing ETL keeps calling CatalogBuilder exactly as it does today (it loads its own DataFrames via etl/loaders/*.py and never touches this class) — this strategy exists for a NEW caller that wants build_from_strategy to also do the Excel/CSV loading step.

Methods

wittgenstein_catalog_builder.sources.ExcelCsvSourceStrategy.extract

method ExcelCsvSourceStrategy.extract() → pd.DataFrame

wittgenstein_catalog_builder.sources.SapHanaSourceStrategy

class SapHanaSourceStrategy(*, host: str, port: int, user: str, password: str, schema: str, package: str, view: str, encrypt: bool = True)

Bases : SourceStrategy

Reads one purpose-built SAP HANA calculation view via hdbcli.

Optional dependency (the sap extra): hdbcli is imported lazily so installing this library for the Excel/CSV path (the original use) never requires it.

Attributes

  • qualified_view_name : str — The double-quoted HANA identifier for this view.

Methods

wittgenstein_catalog_builder.sources.SapHanaSourceStrategy.qualified_view_name

property SapHanaSourceStrategy.qualified_view_name: str

The double-quoted HANA identifier for this view.

Both segments need quoting: the schema (_SYS_BIC, by SAP convention) and <package>/<view> because of the literal / a calculation view's technical name carries.

wittgenstein_catalog_builder.sources.SapHanaSourceStrategy.extract

method SapHanaSourceStrategy.extract() → pd.DataFrame

wittgenstein_catalog_builder.sources.SapHanaSourceStrategy.extract_chunks

method SapHanaSourceStrategy.extract_chunks(chunk_size: int) → Iterator[pd.DataFrame]