Dataframe Backends#
call-report does not have its own dataframe type. Every reader returns a
native frame of whichever library you have configured, reached through
narwhals so the parsing code
is written once rather than three times.
Three backends are supported: pandas (the default), polars, and
pyarrow. They are optional dependencies, so install the one you want:
pip install "call-report[polars]"
Choosing a backend#
get_config() reports the current settings:
>>> from call_report.config import config_context, get_config, set_config
>>> get_config()
{'dataframe_backend': 'pandas', 'lazy': False}
config_context() changes them for the duration of
a with block and restores them on exit, which is the safer choice
whenever the change is meant to be temporary:
>>> from call_report.fca import FCACallReport
>>> from call_report.fca.transport import PackagedArchiveTransport
>>> report = FCACallReport(
... start="2025-03-31", end="2025-03-31", transport=PackagedArchiveTransport()
... )
>>> with config_context(dataframe_backend="pyarrow"):
... type(report.load(schedule="RCF1")).__name__
...
'Table'
set_config() changes them for the rest of the
process:
>>> set_config(dataframe_backend="polars")
>>> type(report.load(schedule="RCF1")).__name__
'DataFrame'
>>> set_config(dataframe_backend="pandas")
Lazy frames#
The lazy setting asks for a frame that defers its work until collected.
Only polars supports it, and asking for it with another backend raises
rather than silently returning an eager frame:
>>> with config_context(dataframe_backend="polars", lazy=True):
... frame = report.load(schedule="RCF1")
...
>>> type(frame).__name__
'LazyFrame'
>>> frame.collect().shape
(767, 14)
Asking for one type at a time#
Every reader also takes a dataframe_type argument, converting its result
as a final step. Use it when the code consuming the frame needs a specific
type while the package stays configured for another:
>>> type(report.load(schedule="RCF1", dataframe_type="pyarrow_table")).__name__
'Table'
The accepted values are "pandas", "pyarrow_table",
"polars_dataframe", and "polars_lazyframe".
What dtypes you get#
Column types come from the schedule’s layout, not from whatever a particular quarter’s values happen to look like. A field that is empty one quarter and populated the next therefore keeps one type across both.
Under pandas that means the nullable extension dtypes rather than the
numpy-backed defaults, since a numpy int64 column cannot hold a missing
value:
>>> rcf1 = report.load(schedule="RCF1")
>>> rcf1["LOANSTATUS"].dtype
Int64Dtype()
That matches what the schedule’s canonical metadata says the field is, so Schemas and Schedule Metadata can be trusted as a description of the frames you actually get:
>>> report.get_file_metadata(schedule="RCF1").file_schema.schema["LOANSTATUS"]
Int64
The period column#
The period column every reader adds names a calendar quarter end.
polars and pyarrow hold it as a date:
>>> with config_context(dataframe_backend="polars"):
... report.load(schedule="RCF1").schema["period"]
...
Date
pandas has no date dtype, so it holds the same value as a datetime whose
time component is always midnight. That keeps the .dt accessor and a
parquet round trip working:
>>> rcf1["period"].dtype
dtype('<M8[us]')
A dataframe_type override follows the same rule, so the dtype depends on
the type you ask for rather than on the backend that built the frame.
Next steps#
Reshaping Across Schedules stacks every schedule into one wide or long frame.
Schemas and Schedule Metadata covers what each column means and how a schedule’s fields have changed over time.