Everything in a dashed
box — every number, table and figure — was captured by
docs/src/build_tutorial.py
actually executing conscipy 0.1.0. None of it is typed in by hand. The Orange
screenshots are real widgets with real data sent through them. You should get
the same numbers — the synthetic data is seeded.
Getting started
1What this is
Python's air-physics layer is not short of options: psychrolib,
CoolProp, psychrochart and MetPy are all
mature. What is missing is the conservation layer — preservation
index, lifetime multiplier, Zeng mould isopleths, the VTT mould index, wood
equilibrium moisture content, and the Getty classification that says whether a
reading needs heating, humidification or dehumidification. That gap is what this
fills.
It is two packages:
| Package | What it is | Depends on |
|---|---|---|
conscipy | The calculation library. Pure NumPy/pandas. | numpy, pandas (matplotlib for plots) |
Orange3-Conservation | Five Orange widgets built on it. | Orange3, conscipy, pyqtgraph |
Both are Python ports of the R package
ConSciR v0.3.0, under GPL-3.
Function names deliberately match the R originals (calcDP,
calcPI, …), so the R documentation and existing scripts transfer
across directly.
Jupyter / local: the full thing. Plots, and room to write your
own analysis.
Google Colab: nothing to install, good for a first look or a
teaching session. conscipy only — Orange is a Qt desktop application and Colab
has no display.
Orange Data Mining: no code at all. Drag widgets onto a canvas
and wire them up.
2Installing
2.1 Local (Jupyter or plain Python)
Neither package is on PyPI yet, so install from GitHub. Make a virtual
environment first — Homebrew Python on macOS and recent Linux distributions will
refuse outright with externally-managed-environment (PEP 668). That
is a guard rail, not a fault.
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS / Linux:
source .venv/bin/activate
pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
pip install matplotlib jupyterlab
On macOS use python3; there is no bare python
command.
2.2 Google Colab
Open a new notebook, paste this into the first cell and press Shift+Enter:
!pip install -q "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
matplotlib and pandas ship with Colab, so there is nothing else to install.
Then import conscipy as cp and you are away.
There is a companion notebook for this tutorial:
Orange3-Conservation is a Qt desktop GUI and needs a display. Colab is
headless: pip install will succeed and the widgets will never open.
Colab covers the conscipy half only.
2.3 Orange Data Mining
Install Orange itself from
orangedatamining.com/download.
The step that matters is the next one: the add-on has to go into the
Python that is running Orange, not the one on your $PATH.
The safest route is Orange's own dialog — Options ▸ Add-ons ▸ Add more… — which always installs into the right place. That needs the package to be on PyPI, though, so until it is published use the terminal.
Windows (using the python in Orange's install directory):
pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=orange3-conservation"
macOS, if Orange came from the .dmg, its Python lives inside the application bundle:
/Applications/Orange3.app/Contents/MacOS/pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
/Applications/Orange3.app/Contents/MacOS/pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=orange3-conservation"
Orange's macOS launcher sets PYTHONNOUSERSITE=1 and
PYTHONSAFEPATH=1. Anything installed into the system Python — or with
pip install --user — is invisible to it, however
cheerful the install log was. The symptom is restarting Orange and finding no
Conservation Science category in the widget toolbox.
Restart Orange after installing. A Conservation Science category appears in the toolbox sidebar.
3Your first calculation
Check the install took. These three numbers come from the ConSciR README and are the reference values the port is checked against (section 21).
import conscipy as cp
print("conscipy", cp.__version__)
print("calcDP(21.8, 36.8) =", cp.calcDP(21.8, 36.8))
print("calcPI(21.8, 36.8) =", cp.calcPI(21.8, 36.8))
print("calcRH_AH(23.8, 7.0524)=", cp.calcRH_AH(23.8, 7.0524))
actual output
conscipy 0.1.0 calcDP(21.8, 36.8) = 6.383969538523943 calcPI(21.8, 36.8) = 45.25849056382369 calcRH_AH(23.8, 7.0524)= 32.819637454815386
Reading those: air at 21.8 °C / 36.8 %RH has a dew point of 6.38 °C and a preservation index of 45.3 years; heat the same air to 23.8 °C without changing its moisture content (7.0524 g/m³) and the relative humidity falls to 32.8 %.
That last one is among the most common questions in the job — what happens to the humidity when the heating goes on in winter?
Data
4Where the data comes from
The short answer: ConSciR's own mydata is usable,
and everything from here on uses it. It is that package's demonstration
dataset, GPL-3 along with the rest of it (DESCRIPTION says
License: GPL (>= 3), with no separate licence on the data), so it
may be used with attribution.
There is no need to copy 3 MB of someone else's data into your own repository, though. Read it from the upstream URL, which works on Colab too:
URL = ("https://raw.githubusercontent.com/BhavShah01/ConSciR/"
"master/inst/extdata/mydata.xlsx")
real = cp.tidy_TRHdata(URL, sheet="mydata")
print(real.head())
actual output
Site Sensor Date Temp RH 0 London External 2024-01-01 00:00:00.000 7.4 79.1 1 London External 2024-01-01 00:15:00.000 7.3 79.5 2 London External 2024-01-01 00:29:59.990 7.2 79.2 3 London External 2024-01-01 00:44:59.985 6.8 80.2 4 London External 2024-01-01 00:59:59.980 6.8 80.8 shape: (104826, 5) site: ['London'] sensors: ['External', 'Room 1', 'Room 1 Case'] date range: 2024-01-01 00:00:00 -> 2024-12-31 23:44:59.610000 interval: 0 days 00:15:00 the file holds 105408 rows; 582 have no reading at all, and tidy_TRHdata drops them, leaving 104826.
Three sensors, which is the point
It holds outdoors, indoors and a display case — a full year at fifteen-minute resolution. That trio demonstrates the central idea of the whole discipline in one table: buffering.
actual output
Temp RH
min mean max min mean max
Sensor
External 0.7 13.4 32.6 20.4 68.5 96.8
Room 1 16.0 20.9 26.2 20.5 46.1 76.0
Room 1 Case 16.3 21.7 25.9 27.0 42.2 57.9
annual range (max - min):
Temp RH
Sensor
External 31.9 76.4
Room 1 10.2 55.5
Room 1 Case 9.6 30.9
Outdoors the year spans 31.9 °C and 76.4 percentage points of humidity. Inside the building that collapses to 10.2 and 55.5; inside the case, to 9.6 and 30.9. Each enclosure absorbs a further share of what the outside is doing — which is what a display case is for, and the numbers say it without help.
Upstream documents this only as @source Climate and "a climate
dataset for use to demonstrate how the functions work".
Whether it is measured data from a real building cannot be
confirmed — the 582 gaps are suggestive of a real logger, but that is
inference. Treating it as a dataset with realistic building behaviour is
safe; citing it as one institution's measurements is not.
So what is example_data() still for?
The package also ships a synthetic generator. It does not exist because the upstream data is unavailable — an earlier version of this page said so, which was wrong, and has been corrected. It exists because a generator gives things a fixed dataset cannot:
df = cp.example_data(n_days=90) # n_days, freq, seed
actual output
Site Sensor Date Temp RH 0 London Room 1 2024-01-01 00:00:00 12.1 57.5 1 London Room 1 2024-01-01 00:15:00 12.0 52.3 2 London Room 1 2024-01-01 00:30:00 12.2 56.3 shape: (17280, 5) sensors: ['Room 1', 'Room 2']
| Why a generator | What a fixed dataset cannot do |
|---|---|
| Seeded | Tests need a deterministic result |
| No network, no file | CI and offline work |
| Kilobytes, not megabytes | Package size |
| Any conditions you like | A real building will not break on cue (section 14) |
a = cp.example_data(n_days=2, seed=0)
b = cp.example_data(n_days=2, seed=0)
c = cp.example_data(n_days=2, seed=1)
print("seed 0 twice identical :", a.equals(b))
print("seed 0 vs seed 1 equal :", a.equals(c))
actual output
seed 0 twice identical : True seed 0 vs seed 1 equal : False
It can show the workflow runs, that shapes are right, that zones get assigned. It cannot show the formulae are correct — that comes from checking against upstream values and from internal consistency. Section 20 draws the line.
5Reading your own logger files
For real work the data comes off a Meaco, a Hanwell, or something else.
tidy_TRHdata() maps the common column-name variants onto the standard
set (Site / Sensor / Date / Temp / RH) and also accepts a file path, CSV or
Excel.
raw = pd.DataFrame({
"Date/Time": ["2024-03-01 09:00", "2024-03-01 09:07", "2024-03-01 09:15"],
"Temperature (C)": [18.4, 18.5, 18.6],
"RH (%)": [52.1, 51.8, 51.5],
"Room": ["Gallery 1"] * 3,
})
tidy = cp.tidy_TRHdata(raw, round_to="15min") # or cp.tidy_TRHdata("logger.csv")
actual output
raw columns: ['Date/Time', 'Temperature (C)', 'RH (%)', 'Room'] Site Sensor Date Temp RH 0 <NA> Gallery 1 2024-03-01 09:00:00 18.4 52.1 1 <NA> Gallery 1 2024-03-01 09:00:00 18.5 51.8 2 <NA> Gallery 1 2024-03-01 09:15:00 18.6 51.5 dtypes: Site object Sensor object Date datetime64[ns] Temp float64 RH float64 dtype: object
Note that Room was recognised as Sensor,
Date/Time as Date and converted to a datetime, and that
round_to="15min" snapped the 09:07 reading to 09:15. Three
readers:
| Function | For |
|---|---|
tidy_TRHdata | Generic / auto. Tables that are already long-format. |
tidy_Meaco | Meaco exports: one Date column plus paired temperature/humidity columns per sensor, reshaped to long. |
tidy_Hanwell | Hanwell exports: a metadata preamble before the data block, whose start is found automatically. |
What a real vendor export looks like
That upstream workbook carries Meaco, Hanwell and
VAISALA sheets alongside mydata — precisely what
these readers exist for. Run them on the real thing:
cp.tidy_Meaco(URL, sheet="Meaco")
cp.tidy_Hanwell(URL, sheet="Hanwell")
actual output
--- tidy_Meaco: 10 readings --- Site Sensor Date Temp RH 0 York Store 2024-01-01 00:03:00 18.17 54.61 1 York Store 2024-01-01 00:13:00 18.16 54.76 site ['York'] sensor ['Store'] 2024-01-01 00:03:00 -> 2024-01-01 01:38:00 --- tidy_Hanwell: 7075 readings --- Site Sensor Date Temp RH 0 None External 2025-08-12 00:06:03 22.8 60.2 1 None External 2025-08-12 00:11:01 22.6 60.2 site [] sensor ['External'] 2025-08-12 00:06:03 -> 2025-09-10 10:30:19
The two formats differ a lot. Meaco is long — one row per reading, columns
named RECEIVER / TRANSMITTER / DATE /
TEMPERATURE / HUMIDITY — so the reader only renames.
Hanwell has a metadata preamble, the sensor name on its own line, the data
block, and then a statistics summary followed by an
---- END OF DATA ---- marker.
mydata is London across 2024; the Meaco sheet is a store room
in York; the Hanwell sheet's preamble says 2025-08-12 to 2025-09-12. They are
three independent examples of format.
Testing them against the sheets above — while writing this page — found
that _read_any() passed header="infer" to
pd.read_excel(), a value only read_csv accepts, so
every Excel path raised ValueError; and that
tidy_Hanwell() mistook the preamble's "Date Start:" line for the
header. Both are fixed and covered by tests. The lesson generalises: without
a real file, "implemented" and "works" are different claims.
The calculations
6Scalars and arrays
Every calc* function is NumPy-vectorised: give it a scalar and you
get a float, give it an array or Series and you get the same shape back. Units
follow the R original — temperature in °C, relative humidity in percent (0–100),
pressure in hPa.
import numpy as np
temps = np.array([15.0, 20.0, 25.0, 30.0])
print("calcPws :", np.round(cp.calcPws(temps), 4))
print("calcDP @55%:", np.round(cp.calcDP(temps, 55.0), 4))
actual output
input Temp : [15. 20. 25. 30.] calcPws : [17.0517 23.3834 31.6853 42.4513] calcDP @55%: [ 6.0302 10.6855 15.3345 19.9773]
| Function | Returns | Unit |
|---|---|---|
calcPws(T) | Saturation vapour pressure | hPa |
calcPw(T, RH) | Partial vapour pressure | hPa |
calcDP(T, RH) | Dew point | °C |
calcFP(T, RH) | Frost point | °C |
calcAH(T, RH) | Absolute humidity | g/m³ |
calcAD(T, RH) | Moist air density | kg/m³ |
calcMR(T, RH) / calcHR | Mixing ratio / humidity ratio | g/kg dry air |
calcSH(T, RH) | Specific humidity | g/kg moist air |
calcEnthalpy(T, RH) | Enthalpy | kJ/kg |
calcFtoC / calcCtoF | Fahrenheit ↔ Celsius | ° |
7Four vapour pressure formulae
calcPws offers four. Which to pick? Across the temperature range
conservation work cares about, the difference is immaterial — but it is better to
know by how much:
for m in ("Buck", "IAPWS", "Magnus", "VAISALA"):
print(f"{m:<8} {cp.calcPws(20.0, method=m):.4f} hPa")
actual output
Buck 23.3834 hPa IAPWS 23.3919 hPa Magnus 23.3344 hPa VAISALA 23.3925 hPa spread 0.0581 hPa (0.248 % of the mean)
They span about 0.25 %, far below the measurement uncertainty of any logger
(typically ±2–3 %RH). The default, Buck, switches to ice coefficients
below 0 °C automatically; leave it alone unless you have a reason.
8The inverse functions
What conservation work actually asks is usually the inverse question: how warm does this air need to be to sit at 50 %RH? Three functions answer it — and the fact that they round-trip is itself a check.
T, RH = 21.8, 36.8
dp = cp.calcDP(T, RH)
ah = cp.calcAH(T, RH)
print(f" RH back from DP {cp.calcRH_DP(T, dp):.6f} %")
print(f" RH back from AH {cp.calcRH_AH(T, ah):.6f} %")
print(f" T needed for {RH} % at that DP: {cp.calcTemp(RH, dp):.6f} degC")
print(f" same, Buck bisection: {cp.calcTemp(RH, dp, method='Buck'):.6f} degC")
actual output
T=21.8 degC, RH=36.8 % dew point 6.383970 degC RH back from DP 36.800000 % (should be 36.8) absolute humidity 7.052415 g/m3 RH back from AH 36.800000 % (should be 36.8) T needed for 36.8 % at that DP: 21.800000 degC same, Buck bisection: 21.794722 degC
calcTemp has two solvers. Magnus has a closed form.
Buck does not: the R version solves it point by point with scalar
uniroot(), while this port uses vectorised bisection,
so it takes whole arrays. The two agree to four decimal places above; the
remainder is the bisection tolerance.
DataFrame workflow
9Humidity columns
In practice you do not compute one value at a time — you hand over a table and
get columns back. That is what the add_* verbs are for.
out = cp.add_humidity_calcs(df.head(3))
print(out.round(3).T)
actual output (transposed for readability)
0 1 2 Site London London London Sensor Room 1 Room 1 Room 1 Date 2024-01-01 00:00:00 2024-01-01 00:15:00 2024-01-01 00:29:59.990000 Temp 21.8 21.8 21.8 RH 36.8 36.7 36.6 Pws 26.121 26.121 26.121 Pw 9.613 9.586 9.56 DP 6.384 6.344 6.305 AH 7.052 7.033 7.014 AD 1.192 1.192 1.192 MR 5.957 5.941 5.925 SH 0.856 0.856 0.856 Enthalpy 37.157 37.115 37.074
Eight columns in one call: Pws, Pw, DP, AH, AD, MR, SH, Enthalpy. Use
method= to change the saturation vapour pressure formula and
P_atm= for atmospheric pressure, which matters at altitude.
10Conservation risk columns
year = cp.example_data(n_days=365)
out = cp.add_conservation_calcs(year)
print(out[["Temp","RH","Mould_LIM","Mould_risk","Mould_rate",
"Mould_index","PreservationIndex","Lifetime","EMC_wood"]].head(3).round(3).T)
actual output
0 1 2 Sensor External External External Temp 7.4 7.3 7.2 RH 79.1 79.5 79.2 Mould_LIM 84.52 84.649 84.779 Mould_risk No risk No risk No risk Mould_rate 0.0 0.0 0.0 Mould_index 0.0 0.0 0.0 PreservationIndex 139.388 140.614 143.11 Lifetime 0.857 0.855 0.856 EMC_wood 16.128 16.269 16.167 readings above the Zeng isopleth, per sensor: External 8497 / 34825 (24.4 %) Room 1 4 / 35127 (0.0 %) Room 1 Case 0 / 34874 (0.0 %) peak VTT mould index, per sensor: Sensor External 1.4668 Room 1 0.0052 Room 1 Case 0.0015
| Column | Meaning |
|---|---|
Mould_LIM | Zeng isopleth: the critical relative humidity (%) at which growth begins, at this temperature. |
Mould_risk | Flagged when RH > Mould_LIM. |
Mould_rate | The growth rate this reading implies, mm/day. |
Mould_index | VTT (Viitanen) mould index, 0–6. Cumulative, not instantaneous. |
PreservationIndex | Years until noticeable deterioration. |
Lifetime | Lifetime multiplier, 1.0 at 20 °C / 50 %RH. |
EMC_wood | Equilibrium moisture content of wood (%). |
The three sensors separate cleanly, and in the right direction: outdoors nearly a quarter of readings sit above the Zeng isopleth and the VTT index climbs to 1.47; indoors only four readings cross; inside the case, none. That is what a building behaving normally looks like.
It is not cp.add_conservation_calcs(real) in one call — it
groups by sensor first. Calling it on the whole table gives a
wrong answer, for the reason in section 16.
11Envelope and adjustments
This is the core of the Getty Tools for the Analysis of Collection Environments framework: given a target temperature and humidity envelope, classify every reading and quantify what it would take to bring it inside.
out = cp.add_humidity_adjustments(year, LowT=16, HighT=25, LowRH=40, HighRH=60)
print(out["zone"].value_counts())
print((out["zone"].value_counts() / 4).round(1)) # readings are 15 min apart -> hours
actual output
Room 1, hours per zone over the year: zone Within 5502 Hum or cooling 2190 Dehum or heating 552 Hum only 392 Dehum only 70 Cooling and dehum 54 Cooling only 21 first three readings needing action: Temp RH zone dTemp_TRHadj dRH_TRHadj 0 21.8 36.8 Hum or cooling 0.0 3.2 1 21.8 36.7 Hum or cooling 0.0 3.3 2 21.8 36.6 Hum or cooling 0.0 3.4
zone takes eleven values, each naming an action:
Within, Heating only, Dehum or heating,
Dehum only, Hum only, Cooling and dehum,
Heating and hum, Heating and dehum,
Hum or cooling, Cooling only,
Cooling and hum.
The "or" zones (Dehum or heating, Hum or cooling) are
meaningful: either intervention would get you there, so you can
choose whichever costs less energy. That is why three sets of adjustment columns
are computed:
| Suffix | Strategy | Columns |
|---|---|---|
_TRHadj | Adjust both temperature and humidity | newTemp / newRH / newAH / dTemp / dRH / dAH |
_AHadj | Absolute humidity only (humidify / dehumidify) | newAH / newRH / dAH / dRH |
_Tadj | Temperature only (heat / cool) | newTemp / newRH / newAH / dTemp / dRH / dAH |
Feed dTemp_TRHadj and dAH_TRHadj into the
HVAC functions and the zone classification turns into
kilowatt-hours.
12Time variables
out = cp.add_time_vars(df.head(2))
print(out[["Date","day","hour","weekday","month","year","Summer"]])
actual output
Date day hour weekday month year Summer 0 2024-01-01 00:00:00 2024-01-01 0 Mon Jan 2024 Winter 1 2024-01-01 00:15:00 2024-01-01 0 Mon Jan 2024 Winter
Summer and DayYear depend on when you run it
add_time_vars uses pd.Timestamp.now().year as its
reference year, projecting each reading's month and day into the current year
before deciding the season. Those two columns are therefore a function of the
system clock as well as the data — worth knowing if you need a reproducible
report.
13Plots
Three plotting functions, each returning a matplotlib Figure you
can keep editing or save. Needs pip install matplotlib.
fig = cp.graph_TRH(year) # T/RH time series, one panel per Sensor
fig = cp.graph_psychrometric(year) # psychrometric chart
fig = cp.graph_TRHbivariate(year) # bivariate density
fig.savefig("chart.png", dpi=150)
graph_TRH(real) — red is temperature, blue relative
humidity, and the two pale bands are the target envelope (16–25 °C,
40–60 %RH by default). The three panels are outdoors, indoors and the
case: the swing being absorbed layer by layer, more legible here than in
section 4's table.graph_psychrometric(real) — grey lines are relative
humidity isolines from 10 % to 100 %, the green block is the target
envelope. The three clouds separate plainly: the outdoor one long and diffuse,
the indoor and case clouds tight around the envelope.graph_TRHbivariate(room) — Room 1 as a bivariate
density, with the envelope as a dashed white box. Easier than a scatter plot
for seeing where the time is actually spent.The chart's vertical axis can be any calculated quantity, via
y_func=:
fig = cp.graph_psychrometric(year, y_func="calcPI") # y axis: preservation index
cp.Y_FUNCS; you can also pass any
callable taking (Temp, RH).Advanced
14Two mould models
There are two independent mould models in the package, answering different questions.
The Zeng isopleth — "is this moment dangerous?"
calcMould_Zeng is memoryless: it looks only at the
current temperature and humidity and answers "above what relative humidity does
growth start, at this temperature?"
for lim in (0, 0.1, 0.5, 1, 2, 3, 4):
print(f"{lim} mm/day 20 degC: {cp.calcMould_Zeng(20.0, 50.0, LIM=lim):.2f}")
# label=True inverts it: give a real reading, get the implied growth rate
cp.calcMould_Zeng(25.0, 90.0, label=True)
actual output
critical RH (%) at which each growth rate starts:
0 mm/day 20 degC: 75.62 25 degC: 74.50
0.1 mm/day 20 degC: 77.09 25 degC: 75.97
0.5 mm/day 20 degC: 80.84 25 degC: 79.72
1 mm/day 20 degC: 83.59 25 degC: 82.47
2 mm/day 20 degC: 86.59 25 degC: 85.47
3 mm/day 20 degC: 89.11 25 degC: 87.99
4 mm/day 20 degC: 91.10 25 degC: 89.98
label=True gives the rate implied by an actual reading:
20.0 degC / 70.0 %RH -> 0.0 mm/day
20.0 degC / 85.0 %RH -> 1.0 mm/day
25.0 degC / 90.0 %RH -> 5.0 mm/day
28.0 degC / 95.0 %RH -> 5.0 mm/day
The VTT index — "how much has accumulated?"
The VTT (Viitanen) model is a recursion: each step's index depends on the one before. It runs from 0 to 6 (0 = no growth, 6 = full coverage), climbing in damp conditions and falling back in dry ones. That memory is the essential difference from Zeng.
Is synthetic data still useful here? Yes, in a different role
Real data tells you what the building did. It cannot tell you what the building might do. To ask what three months of a broken dehumidifier would look like, you have to construct that — and it can be derived from the same real readings rather than from a second generator:
damp = room.copy()
damp["RH"] = np.clip(damp["RH"] + 25.0, 5, 99) # same room, 25 points damper
actual output
Real data covers what the building did, not what it might do. A generator reaches the rest -- here the same room, 25 %RH damper: Room 1 as measured RH max 76.0 % at risk 4 peak index 0.0052 Room 1 + 25 %RH RH max 99.0 % at risk 11415 peak index 1.3293
Same room, same year of temperatures, humidity lifted by 25 points: readings above the isopleth go from 4 to 11,415, and the VTT index from 0.005 to 1.33. Real data cannot do this — it has exactly one outcome.
The four material sensitivity classes
dt = cp.dt_days_from_times(outside["Date"])
for s in cp.VTT_SENSITIVITY: # very / sensitive / medium / resistant
m = cp.mould_index_series(outside["Temp"], outside["RH"],
dt_days=dt, sensitivity=s)
print(f"{s:<10} -> {m[-1]:.4f}")
actual output
final VTT index after a year outdoors: very (A=1.0, B=7.0, C=2.0) -> 1.4668 sensitive (A=0.3, B=6.0, C=1.0) -> 1.4646 medium (A=0.0, B=5.0, C=1.5) -> 1.4578 resistant (A=0.0, B=3.0, C=1.0) -> 1.3733 the same four classes inside Room 1: very -> 0.0052 sensitive -> 0.0052 medium -> 0.0052 resistant -> 0.0052
15Time, recursion and order
The VTT model is a recursion, and three things follow. The first is the one that will bite.
15.1 How much time a reading stands for
The growth rate is per day — the 7 * in
dM_dt is what converts Viitanen's response time from weeks to
days. So every step of the recursion has to know how long that step was. For
fifteen-minute data each step is 1/96 of a day, not a whole one.
dt = cp.dt_days_from_times(outside["Date"]) # days per reading
m = cp.mould_index_series(outside["Temp"], outside["RH"], dt_days=dt)
actual output
External, 34825 readings over 365 days median step 15 minutes largest gap 49.8 hours final mould index, real elapsed time 1.4668 final mould index, one day per reading 6.0000 <- saturated resampled hourly, same year 1.4701
The difference decides the answer: timed properly the index reaches 1.47; counting each reading as a day saturates at 6.0. Resample the same year to hourly and you get 1.4701 — essentially unchanged. That is the point: the result should depend on how long conditions lasted, not on how often the logger wrote them down.
dt_days is required
It has no default, deliberately: any default silently reinstates the bug for
somebody. add_conservation_calcs() derives it from the
Date column, and the Orange widget from its Time column
setting. dt_days_from_times() also handles interruptions — the
largest gap in this dataset is 49.8 hours, and max_gap_days caps
how much elapsed time a gap may contribute.
15.2 The upstream element-wise problem
The R calcMould_VTT() applies itself row by row through
mapply, with M_prev defaulting to 0 — so
every row assumes there was no mould a moment ago. That is
correct for a single time step only. Over a whole series it badly
underestimates.
15.3 Order
actual output
readings 34825 final index, time order 1.466842 final index, shuffled 1.466282 difference 0.038 % Small, but not zero, and only small while growth stays well below M_max. Once k2 starts biting, when the index rises, order matters: as measured peak 1.4668 readings above 1.0: 5387 shuffled peak 1.4663 readings above 1.0: 10952 upstream element-wise (M_prev = 0 every row): max 0.034676 cumulative recursion: max 1.466842
Shuffle the data and the answer is almost, but not exactly,
the same. An earlier version of this page said the result was order-invariant;
that held only because the demo data kept the index so low that k2
stayed at 1 and each increment was independent of M. Once the
index climbs and k2 starts to bite, order matters.
The practical conclusion is unchanged and stronger: sorting is the caller's job. The Orange widget sorts by its time column; the library does not.
16The grouping trap
add_conservation_calcs() does not group by
sensor. It treats the whole table as one continuous time series and runs the
recursion straight through it. And mydata is exactly "the outdoor
year, then the indoor year, then the case" — so you hit this the very first
time you follow along.
whole = cp.add_conservation_calcs(real) # all three at once
for name, g in real.groupby("Sensor"):
print(name, cp.add_conservation_calcs(g)["Mould_index"].max())
actual output
mydata holds three sensors, one after another: ['External', 'Room 1', 'Room 1 Case'] add_conservation_calcs(real) peak index 1.4736 ...on 'External' alone peak index 1.4668 ...on 'Room 1' alone peak index 0.0052 ...on 'Room 1 Case' alone peak index 0.0015 Split by sensor, and sort, before the recursion runs: grouped peak index 1.4668
Each sensor continues from wherever the last one ended rather than starting at zero. Worse, the clock jumps back a year at each join, and that step is credited with no elapsed time at all — so the wrong answer is not merely inflated, it is meaningless. Sort, then group:
fixed = pd.concat(
[cp.add_conservation_calcs(g)
for _, g in real.sort_values(["Sensor", "Date"]).groupby("Sensor")],
ignore_index=True)
concat rather than groupby(...).apply(...) is
deliberate: the latter needs the include_groups= argument, which
pandas 2.2 deprecates.
The Orange Conservation Risk widget does handle this — it has Time column and Group by settings, sorts and groups before running, then writes the results back to the caller's row positions. Library and widget genuinely differ here; using the library, you supply the grouping yourself.
17HVAC loads
Turning "how much adjustment" into "how much energy".
cp.calcCoolingPower(28, 21, 70, 50, 2.0) # 28°C/70% → 21°C/50%, 2 m³/s
cp.calcSensibleHeating(21, 28, 2.0) # sensible load
cp.calcTotalHeating(21, 28, 50, 70, 2.0) # total load
cp.calcSensibleHeatRatio(21, 28, 50, 70, 2.0) # sensible heat ratio (%)
cp.calcCoolingCapacity(3000) # capacity to remove a 3 kW electrical load
cp.calcOffCoilDewPoint(26, 6, 12) # supply air temperature
actual output
Cool 28 degC/70 %RH outside air to 21 degC/50 %RH at 2 m3/s: cooling power 0.0697 kW sensible heating 16.8223 kW total heating 71.7508 kW sensible heat ratio 23.45 % Cooling capacity for a 3 kW lighting load: 4.3714 kW Off-coil temperature, 26 degC on-coil, 6/12 degC chilled water: 10.7000 degC
calcCoolingPower comes out 1000× too small
cooling power 0.0697 kW and total heating 71.75 kW
above describe the same change of air state, three orders of magnitude apart.
Section 23 works it through.
Orange Data Mining
18The five widgets
Once installed, a Conservation Science category appears in the toolbox sidebar. The screenshots below are real widgets, built and fed 30 days of demo data.
| Widget | In | Out | What it does |
|---|---|---|---|
| Logger File | – | Data | Reads Meaco, Hanwell or generic CSV/Excel into a tidy table. |
| Humidity Calculations | Data | Data | Appends any of Pws, Pw, DP, FP, AH, AD, MR, SH, Enthalpy. |
| Conservation Risk | Data | Data | Preservation index, lifetime multiplier, wood EMC, Zeng isopleth, VTT index. |
| Envelope Zones | Data | Data | Classifies against a target envelope and quantifies three strategies. |
| Psychrometric Chart | Data | Selected Data, Data | Interactive chart; drag a box to send readings downstream. |
example_data(). Format selects between the
three tidy functions.Why results are split two ways
A deliberate design decision: continuous results go into X
(features), categorical results into metas. New numeric
columns are then picked up automatically by PCA, PLS and regression, while zone
and risk labels stay in metas, where they can colour a plot or split a group
without distorting distances and projections.
Every DiscreteVariable is also built with its full list of
possible values, not just those present in the batch at hand. Two files
processed separately therefore share one encoding and can be concatenated.
19Building a workflow
A typical chain:
Logger File ──▶ Humidity Calculations ──▶ Conservation Risk ──▶ Envelope Zones
│
▼
Psychrometric Chart ──▶ Data Table
Because the outputs are ordinary Orange Tables, they feed straight
into anything else on the canvas — Data Table, Distributions, Scatter Plot, PCA,
Hierarchical Clustering. That is the main reason to reach for Orange instead of
writing a script.
A ready-made workflow ships with the package. After installing, open Help ▸ Example Workflows ▸ Conservation Science Examples ▸ collection-environment, tick Use synthetic demo data in the Logger File widget, and the whole chain runs without a file of your own.
Verification
20What synthetic data proves
The distinction that matters:
| Synthetic data can verify | Synthetic data cannot verify |
|---|---|
| The pipeline runs end to end | Whether the formulae are right |
| Output shapes, columns and dtypes | Whether a coefficient was mistranscribed |
Zones get assigned, no stray None | Whether the model suits your materials |
| Order invariance of the recursion | Logger measurement error |
| Boundary and extreme cases, built deliberately | How a real building behaves |
Verifying the formulae needs two other things: checking against known values (section 21) and internal consistency (section 22). All three together are what makes the port trustworthy.
21Against upstream values
Correctness of the port rests on comparing Python results with the numbers
printed in the ConSciR README, to a relative tolerance of 1e-6. That is what
conscipy/tests/test_parity.py does:
actual output
call conscipy ConSciR README rel. diff calcDP(21.8, 36.8) 6.38397 6.38397 7.23e-08 calcAH(21.8, 36.8) 7.05242 7.05242 6.72e-07 calcPI(21.8, 36.8) 45.25849 45.25850 2.08e-07 calcRH_AH(23.8, 7.0524) 32.81964 32.81970 1.91e-06
All four land between 1e-6 and 1e-8 — the residual is the README quoting six significant figures, not a disagreement in the calculation.
python -m pytest conscipy/tests -q
That is 19 tests. The Orange widgets have another 39, written against Orange's
own WidgetTest harness, so they exercise the real signal plumbing
rather than the calculation functions in isolation.
Three of those 19, in
conscipy/tests/test_upstream_quirks.py, are unusual: they pin the
current behaviour of the two findings below
and of the grouping trap. The point is to make the
documentation and the code fail together — if someone corrects
calcCoolingPower to match the physics, the test breaks and points at
the section of this page that then also needs updating, rather than leaving the
page quietly describing behaviour the package no longer has.
22Internal consistency
The third technique needs no external reference at all — only relationships that must hold physically. Run them over the synthetic data; any failure is a bug.
df = cp.add_humidity_calcs(cp.example_data(n_days=30))
T, R = df["Temp"].to_numpy(), df["RH"].to_numpy()
assert (df["DP"] <= T).all() # dew point ≤ dry bulb
assert (df["Pw"] <= df["Pws"]).all() # partial ≤ saturation
assert np.allclose(cp.calcRH_DP(T, df["DP"]), R, atol=1e-6) # inverses round-trip
assert np.allclose(cp.calcRH_AH(T, df["AH"]), R, atol=1e-3)
assert (df["SH"] < df["MR"]).all() # specific < mixing ratio
assert df["AD"].between(1.0, 1.3).all() # plausible air density
actual output
rows checked: 35127 dew point never above dry bulb : True Pw never above Pws : True RH recovered from DP within 1e-6 : True RH recovered from AH within 1e-3 : True specific humidity < mixing ratio : True air density in 1.0-1.3 kg/m3 : True (range 1.1701-1.2173)
These are worth writing as tests precisely because they need no reference values — and because they catch the class of bug where the numbers look entirely reasonable and are wrong. Most of the next section surfaced exactly this way.
But see the limit too. The SH < MR check above passes, and
calcSH is nonetheless wrong by a factor of eight —
an assertion about ordering cannot catch an error of scale.
Section 23.4 has it.
23Five findings
Applying the method above while writing this page turned up five places where results contradict each other or the literature. Four trace to the upstream R package — conscipy is a faithful port, and the NOTICE says the formulae follow the original — and one belongs to this port and has been fixed.
| # | Problem | Origin | Status |
|---|---|---|---|
| 23.1 | calcCoolingPower 1000× low | upstream | ConSciR#4 |
| 23.2 | VTT critical humidity sign | upstream | ConSciR#5 |
| 23.3 | calcLM: direction, units and exponent | upstream | not yet reported |
| 23.4 | calcSH unit conversion | upstream | not yet reported |
| 23.5 | Missing temperature scored as “no risk” | upstream | not yet reported |
| — | VTT recursion had no time base | this port | fixed (15.1) |
23.1 The units in calcCoolingPower
The two numbers in section 17 differ by a factor of 1000. Working it by hand shows why:
actual output
air density rho 1.1605 kg/m3 enthalpy in h1 70.8741 kJ/kg enthalpy out h2 40.8384 kJ/kg rho * V * (h1 - h2) 69.7144 kW <- kg/s * kJ/kg = kW calcCoolingPower(...) 0.0697 kW calcTotalHeating(...) 71.7508 kW ratio between the two 1029.2
calcEnthalpy returns kJ/kg and the mass flow is
kg/s, so their product is already kW. The ConSciR R source reads:
coolingPowerWatts <- massFlowRate * (h1 - h2)
coolingPowerKW <- coolingPowerWatts / 1000
The name coolingPowerWatts gives the assumption away: that
h is in J/kg. calcTotalHeating, in the same package,
has no such division.
What to do: use abs(calcTotalHeating(...)) for
a cooling load.
23.2 The sign in the VTT critical humidity
ConSciR and conscipy compute
-0.00267·T³ + 0.160·T² +
3.13·T + 100, where Hukka & Viitanen (1999) have
− 3.13·T:
actual output
T (degC) ConSciR / conscipy Hukka & Viitanen
5.0 119.32 88.02
10.0 144.63 82.03
15.0 173.94 80.04
19.9 204.61 80.03
20.0 205.24 80.04
20.1 80.00 80.00
25.0 80.00 80.00
A critical RH above 100 % can never be exceeded by a real reading,
so frac = (RHcrit - RH)/(RHcrit - 100) turns positive and M_max grows.
That is why a 12 degC / 53 %RH room accumulates a VTT index at all,
while the Zeng isopleth on the same readings reports no risk.
Look at the 20.0 and 20.1 rows. Above 20 °C the model uses a fixed
80 %; below it, the polynomial. The published coefficients make the
polynomial land on 80.04 at 20 °C, joining
continuously; +3.13T gives 205.24 before dropping
to 80 — a 125-point discontinuity.
The consequence is concrete: when RH_crit exceeds 100 %,
frac = (RH_crit − RH) / (RH_crit − 100) becomes a large
positive number, M_max grows with it, and the model predicts
growth in cold, dry air. Above 20 °C nothing is
affected, and the Zeng isopleth is a separate implementation without the
problem.
23.3 calcLM: three errors, two of which hide each other
The lifetime multiplier is meant to say how long something lasts in these conditions relative to 20 °C / 50 %RH. Michalski's rule of thumb is that 5 °C cooler roughly doubles the lifetime. What comes out instead:
actual output
Lifetime multiplier, as ConSciR and conscipy compute it:
5.0 degC / 50 %RH -> 0.997790
10.0 degC / 50 %RH -> 0.998552
15.0 degC / 50 %RH -> 0.999288
20.0 degC / 50 %RH -> 1.000000
25.0 degC / 50 %RH -> 1.000688
30.0 degC / 50 %RH -> 1.001354
A 5 degC drop should roughly double the lifetime (Michalski);
here 20 -> 15 degC changes it by -0.07 %.
25C/50% 15C/50% 20C/30%
as written EA=100 J - 1/3 1.0007 0.9993 1.1856
units fixed EA=100 kJ - 1/3 1.9899 0.4907 1.1856
units and sign EA=100 kJ + 1/3 0.5025 2.0380 1.1856
published form EA=100 kJ + 1.3 0.5025 2.0380 1.9427
Cooling by 5 °C changes the lifetime by −0.07 %, and in the wrong direction — warmer reads as longer-lived. Three errors compound:
- Units.
EA=100is treated as J/mol where the literature uses 100 kJ/mol, a factor of 1000. The exponential collapses to 1.000x and temperature all but drops out. - Sign. The code has
exp(−EA/R · (1/T − 1/T₀)); the literature has+. - Humidity exponent. The code has
(50/RH)^(1/3); the literature has^1.3.
See the second row of that table. Changing EA=100 to
100e3 — the obvious fix — gives 1.99
at 25 °C: warmer, and twice the lifetime. The unit error was holding
the sign error down at the 0.0007 level where nobody could see it; correcting
the units alone amplifies it a thousandfold. All three have to change
together. They then give 2.04 at 15 °C (5 °C cooler
doubles the lifetime ✓) and 1.94 at 20 °C / 30 %RH
(halving the humidity more than doubles it ✓) — both literature
anchors land.
Until this is corrected upstream, do not read calcLM as
a usable estimate of expected lifetime. calcPI is a
separate formula — IPI's Arrhenius rate — and is unaffected.
23.4 The unit conversion in calcSH
actual output
calcMR(20.0, 50.0) 7.260814 g/kg dry air calcSH(20.0, 50.0) 0.878947 <- MR / (1 + MR) q = w / (1 + w) holds for w in kg/kg. With w in g/kg: w = 7.260814 g/kg = 0.007260814 kg/kg q = 0.007208475 kg/kg = 7.208475 g/kg So the label says g/kg but the value is 8.2x too small. The ordering check SH < MR still passes: True -- which is why it never caught this.
The relation between specific humidity and mixing ratio,
q = w / (1 + w), holds for w in kg/kg. But
calcMR returns g/kg, so substituting directly
loses the factor of 1000. The result is labelled g/kg and is 8.2 times too
small.
The reach is wider than the library: add_humidity_calcs()
writes an SH column every time, Orange's Humidity
Calculations widget offers it as a tick box, and the
Psychrometric Chart will put it on the vertical axis — a
curve labelled g/kg and eight times too low.
One of the internal consistency checks is SH < MR, and it has
always passed — and it always will, because 0.879 genuinely is less than
7.26. An assertion about ordering cannot catch an error of
scale. The check stays in this tutorial rather than being quietly
removed, because nothing explains what verification does and does not buy you
more clearly.
23.5 A missing temperature scores as “no risk”
actual output
A reading with no temperature, at 90 %RH: calcMould_Zeng(nan, 90) nan calcMould_Zeng(nan, 90, label=True) 0.0 <- reads as 'no growth' calcMould_Zeng(20, nan, label=True) nan mydata has 582 rows with no reading. tidy_TRHdata drops those, but a logger that reports humidity while its thermistor is faulty would not be dropped, and would be scored as safe.
calcMould_Zeng(..., label=True) checks whether the humidity is
NaN but not the temperature. With no temperature every comparison is False and
the result stays at 0.0 — and 0.0 means
“no growth”. The DataFrame and the Orange widget then encode that as
No risk, indistinguishable from a genuinely safe
reading.
It should return NaN for the numeric result and a missing category, not No
risk. tidy_TRHdata() drops rows where both readings are absent, so
this only bites when humidity is fine and the thermistor has failed —
which is precisely the case that most needs flagging.
conscipy is positioned as a faithful port of ConSciR, so reproducing
upstream behaviour is deliberate: correcting it would break
parity between the two. conscipy/tests/test_upstream_quirks.py
pins the current behaviour so the documentation and the code fail together.
But it also means this: the package suits teaching and exploration, and
does not yet suit making real collection-risk or HVAC decisions —
not from calcLM, calcSH, or the VTT index below
20 °C.
24References
- Source and issues: github.com/Tai-ShengYeh/conservation-python
- Upstream R package: ConSciR · documentation
- Orange Data Mining: orangedatamining.com
- Cosaert, A., Beltran, V., et al. (2022). Tools for the Analysis of Collection Environments. Getty Conservation Institute.
- Hukka, A. & Viitanen, H. (1999). A mathematical model of mould growth on wooden material. Wood Science and Technology, 33, 475–485. doi:10.1007/s002260050131
If you use this in published work, please cite the original package and the framework it implements:
Shah, B., Cosaert, A., Beltran, V., et al. ConSciR: Tools for Conservation
Science. R package. https://bhavshah01.github.io/ConSciR/