中文English

conservation-python

Preventive conservation in Python: humidity and psychrometric calculations, damage functions, mould models, and the Orange Data Mining widget set that puts all of it on a visual canvas. Installation through to the advanced material and how the port is verified.

conscipy 0.1.0 Orange3-Conservation 0.1.0 Jupyter / Colab / Orange GPL-3.0-or-later
Every output on this page was produced by running the code

Everything in a dashed box — every number, table and figure — was captured by docs/src/build_tutorial.py actually executing conscipy 0.1.0. None of it is typed in by hand. The Orange screenshots are real widgets with real data sent through them. You should get the same numbers — the synthetic data is seeded.

Getting started

1What this is

Python's air-physics layer is not short of options: psychrolib, CoolProp, psychrochart and MetPy are all mature. What is missing is the conservation layer — preservation index, lifetime multiplier, Zeng mould isopleths, the VTT mould index, wood equilibrium moisture content, and the Getty classification that says whether a reading needs heating, humidification or dehumidification. That gap is what this fills.

It is two packages:

PackageWhat it isDepends on
conscipyThe calculation library. Pure NumPy/pandas. numpy, pandas (matplotlib for plots)
Orange3-ConservationFive Orange widgets built on it. Orange3, conscipy, pyqtgraph

Both are Python ports of the R package ConSciR v0.3.0, under GPL-3. Function names deliberately match the R originals (calcDP, calcPI, …), so the R documentation and existing scripts transfer across directly.

Three routes — pick one

Jupyter / local: the full thing. Plots, and room to write your own analysis.
Google Colab: nothing to install, good for a first look or a teaching session. conscipy only — Orange is a Qt desktop application and Colab has no display.
Orange Data Mining: no code at all. Drag widgets onto a canvas and wire them up.

2Installing

2.1 Local (Jupyter or plain Python)

Neither package is on PyPI yet, so install from GitHub. Make a virtual environment first — Homebrew Python on macOS and recent Linux distributions will refuse outright with externally-managed-environment (PEP 668). That is a guard rail, not a fault.

python -m venv .venv
# Windows:  .venv\Scripts\activate
# macOS / Linux:
source .venv/bin/activate
pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
pip install matplotlib jupyterlab

On macOS use python3; there is no bare python command.

2.2 Google Colab

Open a new notebook, paste this into the first cell and press Shift+Enter:

!pip install -q "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"

matplotlib and pandas ship with Colab, so there is nothing else to install. Then import conscipy as cp and you are away.

There is a companion notebook for this tutorial:

Open In Colab

The Orange widgets cannot run on Colab

Orange3-Conservation is a Qt desktop GUI and needs a display. Colab is headless: pip install will succeed and the widgets will never open. Colab covers the conscipy half only.

2.3 Orange Data Mining

Install Orange itself from orangedatamining.com/download. The step that matters is the next one: the add-on has to go into the Python that is running Orange, not the one on your $PATH.

The safest route is Orange's own dialog — Options ▸ Add-ons ▸ Add more… — which always installs into the right place. That needs the package to be on PyPI, though, so until it is published use the terminal.

Windows (using the python in Orange's install directory):

pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=orange3-conservation"

macOS, if Orange came from the .dmg, its Python lives inside the application bundle:

/Applications/Orange3.app/Contents/MacOS/pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=conscipy"
/Applications/Orange3.app/Contents/MacOS/pip install "git+https://github.com/Tai-ShengYeh/conservation-python.git#subdirectory=orange3-conservation"
Installed it and Orange still cannot see it? Almost always the wrong Python

Orange's macOS launcher sets PYTHONNOUSERSITE=1 and PYTHONSAFEPATH=1. Anything installed into the system Python — or with pip install --user — is invisible to it, however cheerful the install log was. The symptom is restarting Orange and finding no Conservation Science category in the widget toolbox.

Restart Orange after installing. A Conservation Science category appears in the toolbox sidebar.

3Your first calculation

Check the install took. These three numbers come from the ConSciR README and are the reference values the port is checked against (section 21).

import conscipy as cp

print("conscipy", cp.__version__)
print("calcDP(21.8, 36.8)     =", cp.calcDP(21.8, 36.8))
print("calcPI(21.8, 36.8)     =", cp.calcPI(21.8, 36.8))
print("calcRH_AH(23.8, 7.0524)=", cp.calcRH_AH(23.8, 7.0524))

actual output

conscipy 0.1.0
calcDP(21.8, 36.8)     = 6.383969538523943
calcPI(21.8, 36.8)     = 45.25849056382369
calcRH_AH(23.8, 7.0524)= 32.819637454815386

Reading those: air at 21.8 °C / 36.8 %RH has a dew point of 6.38 °C and a preservation index of 45.3 years; heat the same air to 23.8 °C without changing its moisture content (7.0524 g/m³) and the relative humidity falls to 32.8 %.

That last one is among the most common questions in the job — what happens to the humidity when the heating goes on in winter?

Data

4Where the data comes from

The short answer: ConSciR's own mydata is usable, and everything from here on uses it. It is that package's demonstration dataset, GPL-3 along with the rest of it (DESCRIPTION says License: GPL (>= 3), with no separate licence on the data), so it may be used with attribution.

There is no need to copy 3 MB of someone else's data into your own repository, though. Read it from the upstream URL, which works on Colab too:

URL = ("https://raw.githubusercontent.com/BhavShah01/ConSciR/"
       "master/inst/extdata/mydata.xlsx")

real = cp.tidy_TRHdata(URL, sheet="mydata")
print(real.head())

actual output

     Site    Sensor                    Date  Temp    RH
0  London  External 2024-01-01 00:00:00.000   7.4  79.1
1  London  External 2024-01-01 00:15:00.000   7.3  79.5
2  London  External 2024-01-01 00:29:59.990   7.2  79.2
3  London  External 2024-01-01 00:44:59.985   6.8  80.2
4  London  External 2024-01-01 00:59:59.980   6.8  80.8

shape:       (104826, 5)
site:        ['London']
sensors:     ['External', 'Room 1', 'Room 1 Case']
date range:  2024-01-01 00:00:00 -> 2024-12-31 23:44:59.610000
interval:    0 days 00:15:00

the file holds 105408 rows; 582 have no
reading at all, and tidy_TRHdata drops them, leaving 104826.

Three sensors, which is the point

It holds outdoors, indoors and a display case — a full year at fifteen-minute resolution. That trio demonstrates the central idea of the whole discipline in one table: buffering.

actual output

             Temp                RH            
              min  mean   max   min  mean   max
Sensor                                         
External      0.7  13.4  32.6  20.4  68.5  96.8
Room 1       16.0  20.9  26.2  20.5  46.1  76.0
Room 1 Case  16.3  21.7  25.9  27.0  42.2  57.9

annual range (max - min):
             Temp    RH
Sensor                 
External     31.9  76.4
Room 1       10.2  55.5
Room 1 Case   9.6  30.9

Outdoors the year spans 31.9 °C and 76.4 percentage points of humidity. Inside the building that collapses to 10.2 and 55.5; inside the case, to 9.6 and 30.9. Each enclosure absorbs a further share of what the outside is doing — which is what a display case is for, and the numbers say it without help.

What cannot be claimed about it

Upstream documents this only as @source Climate and "a climate dataset for use to demonstrate how the functions work". Whether it is measured data from a real building cannot be confirmed — the 582 gaps are suggestive of a real logger, but that is inference. Treating it as a dataset with realistic building behaviour is safe; citing it as one institution's measurements is not.

So what is example_data() still for?

The package also ships a synthetic generator. It does not exist because the upstream data is unavailable — an earlier version of this page said so, which was wrong, and has been corrected. It exists because a generator gives things a fixed dataset cannot:

df = cp.example_data(n_days=90)     # n_days, freq, seed

actual output

     Site  Sensor                Date  Temp    RH
0  London  Room 1 2024-01-01 00:00:00  12.1  57.5
1  London  Room 1 2024-01-01 00:15:00  12.0  52.3
2  London  Room 1 2024-01-01 00:30:00  12.2  56.3

shape: (17280, 5)  sensors: ['Room 1', 'Room 2']
Why a generatorWhat a fixed dataset cannot do
SeededTests need a deterministic result
No network, no fileCI and offline work
Kilobytes, not megabytesPackage size
Any conditions you likeA real building will not break on cue (section 14)
a = cp.example_data(n_days=2, seed=0)
b = cp.example_data(n_days=2, seed=0)
c = cp.example_data(n_days=2, seed=1)
print("seed 0 twice identical :", a.equals(b))
print("seed 0 vs seed 1 equal :", a.equals(c))

actual output

seed 0 twice identical : True
seed 0 vs seed 1 equal : False
Synthetic data verifies the pipeline, not the physics

It can show the workflow runs, that shapes are right, that zones get assigned. It cannot show the formulae are correct — that comes from checking against upstream values and from internal consistency. Section 20 draws the line.

5Reading your own logger files

For real work the data comes off a Meaco, a Hanwell, or something else. tidy_TRHdata() maps the common column-name variants onto the standard set (Site / Sensor / Date / Temp / RH) and also accepts a file path, CSV or Excel.

raw = pd.DataFrame({
    "Date/Time":       ["2024-03-01 09:00", "2024-03-01 09:07", "2024-03-01 09:15"],
    "Temperature (C)": [18.4, 18.5, 18.6],
    "RH (%)":          [52.1, 51.8, 51.5],
    "Room":            ["Gallery 1"] * 3,
})
tidy = cp.tidy_TRHdata(raw, round_to="15min")   # or cp.tidy_TRHdata("logger.csv")

actual output

raw columns: ['Date/Time', 'Temperature (C)', 'RH (%)', 'Room']

   Site     Sensor                Date  Temp    RH
0  <NA>  Gallery 1 2024-03-01 09:00:00  18.4  52.1
1  <NA>  Gallery 1 2024-03-01 09:00:00  18.5  51.8
2  <NA>  Gallery 1 2024-03-01 09:15:00  18.6  51.5

dtypes:
Site              object
Sensor            object
Date      datetime64[ns]
Temp             float64
RH               float64
dtype: object

Note that Room was recognised as Sensor, Date/Time as Date and converted to a datetime, and that round_to="15min" snapped the 09:07 reading to 09:15. Three readers:

FunctionFor
tidy_TRHdataGeneric / auto. Tables that are already long-format.
tidy_MeacoMeaco exports: one Date column plus paired temperature/humidity columns per sensor, reshaped to long.
tidy_HanwellHanwell exports: a metadata preamble before the data block, whose start is found automatically.

What a real vendor export looks like

That upstream workbook carries Meaco, Hanwell and VAISALA sheets alongside mydata — precisely what these readers exist for. Run them on the real thing:

cp.tidy_Meaco(URL, sheet="Meaco")
cp.tidy_Hanwell(URL, sheet="Hanwell")

actual output

--- tidy_Meaco: 10 readings ---
   Site Sensor                Date   Temp     RH
0  York  Store 2024-01-01 00:03:00  18.17  54.61
1  York  Store 2024-01-01 00:13:00  18.16  54.76
  site ['York']  sensor ['Store']
  2024-01-01 00:03:00 -> 2024-01-01 01:38:00

--- tidy_Hanwell: 7075 readings ---
   Site    Sensor                Date  Temp    RH
0  None  External 2025-08-12 00:06:03  22.8  60.2
1  None  External 2025-08-12 00:11:01  22.6  60.2
  site []  sensor ['External']
  2025-08-12 00:06:03 -> 2025-09-10 10:30:19

The two formats differ a lot. Meaco is long — one row per reading, columns named RECEIVER / TRANSMITTER / DATE / TEMPERATURE / HUMIDITY — so the reader only renames. Hanwell has a metadata preamble, the sensor name on its own line, the data block, and then a statistics summary followed by an ---- END OF DATA ---- marker.

These are three separate exports, not one building

mydata is London across 2024; the Meaco sheet is a store room in York; the Hanwell sheet's preamble says 2025-08-12 to 2025-09-12. They are three independent examples of format.

Two of these three readers did not work until recently

Testing them against the sheets above — while writing this page — found that _read_any() passed header="infer" to pd.read_excel(), a value only read_csv accepts, so every Excel path raised ValueError; and that tidy_Hanwell() mistook the preamble's "Date Start:" line for the header. Both are fixed and covered by tests. The lesson generalises: without a real file, "implemented" and "works" are different claims.

The calculations

6Scalars and arrays

Every calc* function is NumPy-vectorised: give it a scalar and you get a float, give it an array or Series and you get the same shape back. Units follow the R original — temperature in °C, relative humidity in percent (0–100), pressure in hPa.

import numpy as np

temps = np.array([15.0, 20.0, 25.0, 30.0])
print("calcPws    :", np.round(cp.calcPws(temps), 4))
print("calcDP @55%:", np.round(cp.calcDP(temps, 55.0), 4))

actual output

input Temp : [15. 20. 25. 30.]
calcPws    : [17.0517 23.3834 31.6853 42.4513]
calcDP @55%: [ 6.0302 10.6855 15.3345 19.9773]
FunctionReturnsUnit
calcPws(T)Saturation vapour pressurehPa
calcPw(T, RH)Partial vapour pressurehPa
calcDP(T, RH)Dew point°C
calcFP(T, RH)Frost point°C
calcAH(T, RH)Absolute humidityg/m³
calcAD(T, RH)Moist air densitykg/m³
calcMR(T, RH) / calcHRMixing ratio / humidity ratiog/kg dry air
calcSH(T, RH)Specific humidityg/kg moist air
calcEnthalpy(T, RH)EnthalpykJ/kg
calcFtoC / calcCtoFFahrenheit ↔ Celsius°

7Four vapour pressure formulae

calcPws offers four. Which to pick? Across the temperature range conservation work cares about, the difference is immaterial — but it is better to know by how much:

for m in ("Buck", "IAPWS", "Magnus", "VAISALA"):
    print(f"{m:<8} {cp.calcPws(20.0, method=m):.4f} hPa")

actual output

Buck     23.3834 hPa
IAPWS    23.3919 hPa
Magnus   23.3344 hPa
VAISALA  23.3925 hPa
spread   0.0581 hPa (0.248 % of the mean)

They span about 0.25 %, far below the measurement uncertainty of any logger (typically ±2–3 %RH). The default, Buck, switches to ice coefficients below 0 °C automatically; leave it alone unless you have a reason.

8The inverse functions

What conservation work actually asks is usually the inverse question: how warm does this air need to be to sit at 50 %RH? Three functions answer it — and the fact that they round-trip is itself a check.

T, RH = 21.8, 36.8
dp = cp.calcDP(T, RH)
ah = cp.calcAH(T, RH)
print(f"  RH back from DP  {cp.calcRH_DP(T, dp):.6f} %")
print(f"  RH back from AH  {cp.calcRH_AH(T, ah):.6f} %")
print(f"  T needed for {RH} % at that DP: {cp.calcTemp(RH, dp):.6f} degC")
print(f"  same, Buck bisection:           {cp.calcTemp(RH, dp, method='Buck'):.6f} degC")

actual output

T=21.8 degC, RH=36.8 %
  dew point           6.383970 degC
  RH back from DP     36.800000 %   (should be 36.8)
  absolute humidity   7.052415 g/m3
  RH back from AH     36.800000 %   (should be 36.8)
  T needed for 36.8 % at that DP: 21.800000 degC
  same, Buck bisection:           21.794722 degC

calcTemp has two solvers. Magnus has a closed form. Buck does not: the R version solves it point by point with scalar uniroot(), while this port uses vectorised bisection, so it takes whole arrays. The two agree to four decimal places above; the remainder is the bisection tolerance.

DataFrame workflow

9Humidity columns

In practice you do not compute one value at a time — you hand over a table and get columns back. That is what the add_* verbs are for.

out = cp.add_humidity_calcs(df.head(3))
print(out.round(3).T)

actual output (transposed for readability)

                            0                    1                           2
Site                   London               London                      London
Sensor                 Room 1               Room 1                      Room 1
Date      2024-01-01 00:00:00  2024-01-01 00:15:00  2024-01-01 00:29:59.990000
Temp                     21.8                 21.8                        21.8
RH                       36.8                 36.7                        36.6
Pws                    26.121               26.121                      26.121
Pw                      9.613                9.586                        9.56
DP                      6.384                6.344                       6.305
AH                      7.052                7.033                       7.014
AD                      1.192                1.192                       1.192
MR                      5.957                5.941                       5.925
SH                      0.856                0.856                       0.856
Enthalpy               37.157               37.115                      37.074

Eight columns in one call: Pws, Pw, DP, AH, AD, MR, SH, Enthalpy. Use method= to change the saturation vapour pressure formula and P_atm= for atmospheric pressure, which matters at altitude.

10Conservation risk columns

year = cp.example_data(n_days=365)
out = cp.add_conservation_calcs(year)
print(out[["Temp","RH","Mould_LIM","Mould_risk","Mould_rate",
           "Mould_index","PreservationIndex","Lifetime","EMC_wood"]].head(3).round(3).T)

actual output

                          0         1         2
Sensor             External  External  External
Temp                    7.4       7.3       7.2
RH                     79.1      79.5      79.2
Mould_LIM             84.52    84.649    84.779
Mould_risk          No risk   No risk   No risk
Mould_rate              0.0       0.0       0.0
Mould_index             0.0       0.0       0.0
PreservationIndex   139.388   140.614    143.11
Lifetime              0.857     0.855     0.856
EMC_wood             16.128    16.269    16.167

readings above the Zeng isopleth, per sensor:
  External       8497 / 34825  (24.4 %)
  Room 1            4 / 35127  (0.0 %)
  Room 1 Case       0 / 34874  (0.0 %)

peak VTT mould index, per sensor:
Sensor
External       1.4668
Room 1         0.0052
Room 1 Case    0.0015
ColumnMeaning
Mould_LIMZeng isopleth: the critical relative humidity (%) at which growth begins, at this temperature.
Mould_riskFlagged when RH > Mould_LIM.
Mould_rateThe growth rate this reading implies, mm/day.
Mould_indexVTT (Viitanen) mould index, 0–6. Cumulative, not instantaneous.
PreservationIndexYears until noticeable deterioration.
LifetimeLifetime multiplier, 1.0 at 20 °C / 50 %RH.
EMC_woodEquilibrium moisture content of wood (%).

The three sensors separate cleanly, and in the right direction: outdoors nearly a quarter of readings sit above the Zeng isopleth and the VTT index climbs to 1.47; indoors only four readings cross; inside the case, none. That is what a building behaving normally looks like.

One detail in that code matters

It is not cp.add_conservation_calcs(real) in one call — it groups by sensor first. Calling it on the whole table gives a wrong answer, for the reason in section 16.

11Envelope and adjustments

This is the core of the Getty Tools for the Analysis of Collection Environments framework: given a target temperature and humidity envelope, classify every reading and quantify what it would take to bring it inside.

out = cp.add_humidity_adjustments(year, LowT=16, HighT=25, LowRH=40, HighRH=60)
print(out["zone"].value_counts())
print((out["zone"].value_counts() / 4).round(1))   # readings are 15 min apart -> hours

actual output

Room 1, hours per zone over the year:
zone
Within               5502
Hum or cooling       2190
Dehum or heating      552
Hum only              392
Dehum only             70
Cooling and dehum      54
Cooling only           21

first three readings needing action:
   Temp    RH            zone  dTemp_TRHadj  dRH_TRHadj
0  21.8  36.8  Hum or cooling           0.0         3.2
1  21.8  36.7  Hum or cooling           0.0         3.3
2  21.8  36.6  Hum or cooling           0.0         3.4

zone takes eleven values, each naming an action: Within, Heating only, Dehum or heating, Dehum only, Hum only, Cooling and dehum, Heating and hum, Heating and dehum, Hum or cooling, Cooling only, Cooling and hum.

The "or" zones (Dehum or heating, Hum or cooling) are meaningful: either intervention would get you there, so you can choose whichever costs less energy. That is why three sets of adjustment columns are computed:

SuffixStrategyColumns
_TRHadjAdjust both temperature and humiditynewTemp / newRH / newAH / dTemp / dRH / dAH
_AHadjAbsolute humidity only (humidify / dehumidify)newAH / newRH / dAH / dRH
_TadjTemperature only (heat / cool)newTemp / newRH / newAH / dTemp / dRH / dAH

Feed dTemp_TRHadj and dAH_TRHadj into the HVAC functions and the zone classification turns into kilowatt-hours.

12Time variables

out = cp.add_time_vars(df.head(2))
print(out[["Date","day","hour","weekday","month","year","Summer"]])

actual output

                 Date        day  hour weekday month  year  Summer
0 2024-01-01 00:00:00 2024-01-01     0     Mon   Jan  2024  Winter
1 2024-01-01 00:15:00 2024-01-01     0     Mon   Jan  2024  Winter
Summer and DayYear depend on when you run it

add_time_vars uses pd.Timestamp.now().year as its reference year, projecting each reading's month and day into the current year before deciding the season. Those two columns are therefore a function of the system clock as well as the data — worth knowing if you need a reproducible report.

13Plots

Three plotting functions, each returning a matplotlib Figure you can keep editing or save. Needs pip install matplotlib.

fig = cp.graph_TRH(year)                    # T/RH time series, one panel per Sensor
fig = cp.graph_psychrometric(year)          # psychrometric chart
fig = cp.graph_TRHbivariate(year)           # bivariate density
fig.savefig("chart.png", dpi=150)
Temperature and humidity time series, one panel per room
graph_TRH(real) — red is temperature, blue relative humidity, and the two pale bands are the target envelope (16–25 °C, 40–60 %RH by default). The three panels are outdoors, indoors and the case: the swing being absorbed layer by layer, more legible here than in section 4's table.
Psychrometric chart with grey RH isolines and a green target envelope
graph_psychrometric(real) — grey lines are relative humidity isolines from 10 % to 100 %, the green block is the target envelope. The three clouds separate plainly: the outdoor one long and diffuse, the indoor and case clouds tight around the envelope.
Two-dimensional histogram of temperature against humidity
graph_TRHbivariate(room) — Room 1 as a bivariate density, with the envelope as a dashed white box. Easier than a scatter plot for seeing where the time is actually spent.

The chart's vertical axis can be any calculated quantity, via y_func=:

fig = cp.graph_psychrometric(year, y_func="calcPI")   # y axis: preservation index
Psychrometric chart with preservation index on the vertical axis
The same room with preservation index (years) on the vertical axis. The available keys are in cp.Y_FUNCS; you can also pass any callable taking (Temp, RH).

Advanced

14Two mould models

There are two independent mould models in the package, answering different questions.

The Zeng isopleth — "is this moment dangerous?"

calcMould_Zeng is memoryless: it looks only at the current temperature and humidity and answers "above what relative humidity does growth start, at this temperature?"

for lim in (0, 0.1, 0.5, 1, 2, 3, 4):
    print(f"{lim} mm/day  20 degC: {cp.calcMould_Zeng(20.0, 50.0, LIM=lim):.2f}")

# label=True inverts it: give a real reading, get the implied growth rate
cp.calcMould_Zeng(25.0, 90.0, label=True)

actual output

critical RH (%) at which each growth rate starts:
     0 mm/day  20 degC:  75.62   25 degC:  74.50
   0.1 mm/day  20 degC:  77.09   25 degC:  75.97
   0.5 mm/day  20 degC:  80.84   25 degC:  79.72
     1 mm/day  20 degC:  83.59   25 degC:  82.47
     2 mm/day  20 degC:  86.59   25 degC:  85.47
     3 mm/day  20 degC:  89.11   25 degC:  87.99
     4 mm/day  20 degC:  91.10   25 degC:  89.98

label=True gives the rate implied by an actual reading:
  20.0 degC / 70.0 %RH -> 0.0 mm/day
  20.0 degC / 85.0 %RH -> 1.0 mm/day
  25.0 degC / 90.0 %RH -> 5.0 mm/day
  28.0 degC / 95.0 %RH -> 5.0 mm/day

The VTT index — "how much has accumulated?"

The VTT (Viitanen) model is a recursion: each step's index depends on the one before. It runs from 0 to 6 (0 = no growth, 6 = full coverage), climbing in damp conditions and falling back in dry ones. That memory is the essential difference from Zeng.

Is synthetic data still useful here? Yes, in a different role

Real data tells you what the building did. It cannot tell you what the building might do. To ask what three months of a broken dehumidifier would look like, you have to construct that — and it can be derived from the same real readings rather than from a second generator:

damp = room.copy()
damp["RH"] = np.clip(damp["RH"] + 25.0, 5, 99)     # same room, 25 points damper

actual output

Real data covers what the building did, not what it might do.
A generator reaches the rest -- here the same room, 25 %RH damper:

  Room 1 as measured     RH max  76.0 %   at risk      4   peak index 0.0052
  Room 1 + 25 %RH        RH max  99.0 %   at risk  11415   peak index 1.3293

Same room, same year of temperatures, humidity lifted by 25 points: readings above the isopleth go from 4 to 11,415, and the VTT index from 0.005 to 1.33. Real data cannot do this — it has exactly one outcome.

The four material sensitivity classes

dt = cp.dt_days_from_times(outside["Date"])
for s in cp.VTT_SENSITIVITY:                 # very / sensitive / medium / resistant
    m = cp.mould_index_series(outside["Temp"], outside["RH"],
                              dt_days=dt, sensitivity=s)
    print(f"{s:<10} -> {m[-1]:.4f}")

actual output

final VTT index after a year outdoors:
  very       (A=1.0, B=7.0, C=2.0)  ->  1.4668
  sensitive  (A=0.3, B=6.0, C=1.0)  ->  1.4646
  medium     (A=0.0, B=5.0, C=1.5)  ->  1.4578
  resistant  (A=0.0, B=3.0, C=1.0)  ->  1.3733

the same four classes inside Room 1:
  very       ->  0.0052
  sensitive  ->  0.0052
  medium     ->  0.0052
  resistant  ->  0.0052
VTT mould index curves for four material sensitivity classes
A year of the outdoor sensor, across the four material sensitivity classes. Flat through winter and spring, climbing through summer and autumn — the model stops growing when it is cold and dry. Note that this curve is affected by the critical-humidity problem in section 23.2, which inflates the low-temperature part.

15Time, recursion and order

The VTT model is a recursion, and three things follow. The first is the one that will bite.

15.1 How much time a reading stands for

The growth rate is per day — the 7 * in dM_dt is what converts Viitanen's response time from weeks to days. So every step of the recursion has to know how long that step was. For fifteen-minute data each step is 1/96 of a day, not a whole one.

dt = cp.dt_days_from_times(outside["Date"])       # days per reading
m  = cp.mould_index_series(outside["Temp"], outside["RH"], dt_days=dt)

actual output

External, 34825 readings over 365 days
  median step   15 minutes
  largest gap   49.8 hours

final mould index, real elapsed time   1.4668
final mould index, one day per reading 6.0000   <- saturated

resampled hourly, same year            1.4701

The difference decides the answer: timed properly the index reaches 1.47; counting each reading as a day saturates at 6.0. Resample the same year to hourly and you get 1.4701 — essentially unchanged. That is the point: the result should depend on how long conditions lasted, not on how often the logger wrote them down.

Which is why dt_days is required

It has no default, deliberately: any default silently reinstates the bug for somebody. add_conservation_calcs() derives it from the Date column, and the Orange widget from its Time column setting. dt_days_from_times() also handles interruptions — the largest gap in this dataset is 49.8 hours, and max_gap_days caps how much elapsed time a gap may contribute.

15.2 The upstream element-wise problem

The R calcMould_VTT() applies itself row by row through mapply, with M_prev defaulting to 0 — so every row assumes there was no mould a moment ago. That is correct for a single time step only. Over a whole series it badly underestimates.

15.3 Order

actual output

readings                   34825
final index, time order    1.466842
final index, shuffled      1.466282
difference                 0.038 %

Small, but not zero, and only small while growth stays well below
M_max. Once k2 starts biting, when the index rises, order matters:
  as measured  peak 1.4668   readings above 1.0: 5387
  shuffled     peak 1.4663   readings above 1.0: 10952

upstream element-wise (M_prev = 0 every row):
  max 0.034676
cumulative recursion:
  max 1.466842

Shuffle the data and the answer is almost, but not exactly, the same. An earlier version of this page said the result was order-invariant; that held only because the demo data kept the index so low that k2 stayed at 1 and each increment was independent of M. Once the index climbs and k2 starts to bite, order matters.

The practical conclusion is unchanged and stronger: sorting is the caller's job. The Orange widget sorts by its time column; the library does not.

16The grouping trap

The easiest mistake to make with this library

add_conservation_calcs() does not group by sensor. It treats the whole table as one continuous time series and runs the recursion straight through it. And mydata is exactly "the outdoor year, then the indoor year, then the case" — so you hit this the very first time you follow along.

whole = cp.add_conservation_calcs(real)                    # all three at once
for name, g in real.groupby("Sensor"):
    print(name, cp.add_conservation_calcs(g)["Mould_index"].max())

actual output

mydata holds three sensors, one after another:
  ['External', 'Room 1', 'Room 1 Case']

add_conservation_calcs(real)      peak index 1.4736
  ...on 'External' alone           peak index 1.4668
  ...on 'Room 1' alone           peak index 0.0052
  ...on 'Room 1 Case' alone           peak index 0.0015

Split by sensor, and sort, before the recursion runs:
  grouped                         peak index 1.4668

Each sensor continues from wherever the last one ended rather than starting at zero. Worse, the clock jumps back a year at each join, and that step is credited with no elapsed time at all — so the wrong answer is not merely inflated, it is meaningless. Sort, then group:

fixed = pd.concat(
    [cp.add_conservation_calcs(g)
     for _, g in real.sort_values(["Sensor", "Date"]).groupby("Sensor")],
    ignore_index=True)

concat rather than groupby(...).apply(...) is deliberate: the latter needs the include_groups= argument, which pandas 2.2 deprecates.

The Orange Conservation Risk widget does handle this — it has Time column and Group by settings, sorts and groups before running, then writes the results back to the caller's row positions. Library and widget genuinely differ here; using the library, you supply the grouping yourself.

17HVAC loads

Turning "how much adjustment" into "how much energy".

cp.calcCoolingPower(28, 21, 70, 50, 2.0)      # 28°C/70% → 21°C/50%, 2 m³/s
cp.calcSensibleHeating(21, 28, 2.0)           # sensible load
cp.calcTotalHeating(21, 28, 50, 70, 2.0)      # total load
cp.calcSensibleHeatRatio(21, 28, 50, 70, 2.0) # sensible heat ratio (%)
cp.calcCoolingCapacity(3000)                  # capacity to remove a 3 kW electrical load
cp.calcOffCoilDewPoint(26, 6, 12)             # supply air temperature

actual output

Cool 28 degC/70 %RH outside air to 21 degC/50 %RH at 2 m3/s:
  cooling power     0.0697 kW
  sensible heating  16.8223 kW
  total heating     71.7508 kW
  sensible heat ratio 23.45 %

Cooling capacity for a 3 kW lighting load: 4.3714 kW
Off-coil temperature, 26 degC on-coil, 6/12 degC chilled water: 10.7000 degC
calcCoolingPower comes out 1000× too small

cooling power 0.0697 kW and total heating 71.75 kW above describe the same change of air state, three orders of magnitude apart. Section 23 works it through.

Orange Data Mining

18The five widgets

Once installed, a Conservation Science category appears in the toolbox sidebar. The screenshots below are real widgets, built and fed 30 days of demo data.

WidgetInOutWhat it does
Logger File–Data Reads Meaco, Hanwell or generic CSV/Excel into a tidy table.
Humidity CalculationsDataData Appends any of Pws, Pw, DP, FP, AH, AD, MR, SH, Enthalpy.
Conservation RiskDataData Preservation index, lifetime multiplier, wood EMC, Zeng isopleth, VTT index.
Envelope ZonesDataData Classifies against a target envelope and quantifies three strategies.
Psychrometric ChartDataSelected Data, Data Interactive chart; drag a box to send readings downstream.
The Logger File widget control panel
Logger File — tick Use synthetic demo data and the whole pipeline runs with no file at all, using example_data(). Format selects between the three tidy functions.
The Humidity Calculations widget control panel
Humidity Calculations — the top two combos map columns; the widget guesses and you can correct it. Both the vapour pressure formula and the atmospheric pressure are adjustable.
The Conservation Risk widget control panel
Conservation Risk — note Time column and Group by: exactly what section 16 says the library lacks. The Info box summarises each run -- readings above the isopleth, and the peak index. It rounds to zero on the demo data because thirty days really is a short exposure: that figure dropped two orders of magnitude when the time base was corrected.
The Envelope Zones widget control panel
Envelope Zones — four envelope bounds plus a choice of adjustment strategy. Output all three strategies emits every column set at once.
The Psychrometric Chart widget with control panel and plot
Psychrometric Chart — interactive. Grey RH isolines, green target envelope, colour by Sensor. Tick Selection box and you can drag a selection; the chosen readings leave through Selected Data.

Why results are split two ways

A deliberate design decision: continuous results go into X (features), categorical results into metas. New numeric columns are then picked up automatically by PCA, PLS and regression, while zone and risk labels stay in metas, where they can colour a plot or split a group without distorting distances and projections.

Every DiscreteVariable is also built with its full list of possible values, not just those present in the batch at hand. Two files processed separately therefore share one encoding and can be concatenated.

19Building a workflow

A typical chain:

Logger File ──▶ Humidity Calculations ──▶ Conservation Risk ──▶ Envelope Zones
                                                    │
                                                    ▼
                                         Psychrometric Chart ──▶ Data Table

Because the outputs are ordinary Orange Tables, they feed straight into anything else on the canvas — Data Table, Distributions, Scatter Plot, PCA, Hierarchical Clustering. That is the main reason to reach for Orange instead of writing a script.

A ready-made workflow ships with the package. After installing, open Help ▸ Example Workflows ▸ Conservation Science Examples ▸ collection-environment, tick Use synthetic demo data in the Logger File widget, and the whole chain runs without a file of your own.

Verification

20What synthetic data proves

The distinction that matters:

Synthetic data can verifySynthetic data cannot verify
The pipeline runs end to endWhether the formulae are right
Output shapes, columns and dtypesWhether a coefficient was mistranscribed
Zones get assigned, no stray NoneWhether the model suits your materials
Order invariance of the recursionLogger measurement error
Boundary and extreme cases, built deliberatelyHow a real building behaves

Verifying the formulae needs two other things: checking against known values (section 21) and internal consistency (section 22). All three together are what makes the port trustworthy.

21Against upstream values

Correctness of the port rests on comparing Python results with the numbers printed in the ConSciR README, to a relative tolerance of 1e-6. That is what conscipy/tests/test_parity.py does:

actual output

call                            conscipy  ConSciR README    rel. diff
calcDP(21.8, 36.8)               6.38397         6.38397     7.23e-08
calcAH(21.8, 36.8)               7.05242         7.05242     6.72e-07
calcPI(21.8, 36.8)              45.25849        45.25850     2.08e-07
calcRH_AH(23.8, 7.0524)         32.81964        32.81970     1.91e-06

All four land between 1e-6 and 1e-8 — the residual is the README quoting six significant figures, not a disagreement in the calculation.

python -m pytest conscipy/tests -q

That is 19 tests. The Orange widgets have another 39, written against Orange's own WidgetTest harness, so they exercise the real signal plumbing rather than the calculation functions in isolation.

Three of those 19, in conscipy/tests/test_upstream_quirks.py, are unusual: they pin the current behaviour of the two findings below and of the grouping trap. The point is to make the documentation and the code fail together — if someone corrects calcCoolingPower to match the physics, the test breaks and points at the section of this page that then also needs updating, rather than leaving the page quietly describing behaviour the package no longer has.

22Internal consistency

The third technique needs no external reference at all — only relationships that must hold physically. Run them over the synthetic data; any failure is a bug.

df = cp.add_humidity_calcs(cp.example_data(n_days=30))
T, R = df["Temp"].to_numpy(), df["RH"].to_numpy()

assert (df["DP"] <= T).all()                                    # dew point ≤ dry bulb
assert (df["Pw"] <= df["Pws"]).all()                            # partial ≤ saturation
assert np.allclose(cp.calcRH_DP(T, df["DP"]), R, atol=1e-6)     # inverses round-trip
assert np.allclose(cp.calcRH_AH(T, df["AH"]), R, atol=1e-3)
assert (df["SH"] < df["MR"]).all()                              # specific < mixing ratio
assert df["AD"].between(1.0, 1.3).all()                         # plausible air density

actual output

rows checked: 35127
dew point never above dry bulb      : True
Pw never above Pws                  : True
RH recovered from DP within 1e-6    : True
RH recovered from AH within 1e-3    : True
specific humidity < mixing ratio    : True
air density in 1.0-1.3 kg/m3        : True  (range 1.1701-1.2173)

These are worth writing as tests precisely because they need no reference values — and because they catch the class of bug where the numbers look entirely reasonable and are wrong. Most of the next section surfaced exactly this way.

But see the limit too. The SH < MR check above passes, and calcSH is nonetheless wrong by a factor of eight — an assertion about ordering cannot catch an error of scale. Section 23.4 has it.

23Five findings

Applying the method above while writing this page turned up five places where results contradict each other or the literature. Four trace to the upstream R package — conscipy is a faithful port, and the NOTICE says the formulae follow the original — and one belongs to this port and has been fixed.

#ProblemOriginStatus
23.1calcCoolingPower 1000× lowupstream ConSciR#4
23.2VTT critical humidity signupstream ConSciR#5
23.3calcLM: direction, units and exponentupstream not yet reported
23.4calcSH unit conversionupstream not yet reported
23.5Missing temperature scored as “no risk”upstream not yet reported
—VTT recursion had no time basethis port fixed (15.1)

23.1 The units in calcCoolingPower

The two numbers in section 17 differ by a factor of 1000. Working it by hand shows why:

actual output

air density rho          1.1605 kg/m3
enthalpy in  h1          70.8741 kJ/kg
enthalpy out h2          40.8384 kJ/kg
rho * V * (h1 - h2)      69.7144 kW   <- kg/s * kJ/kg = kW

calcCoolingPower(...)    0.0697 kW
calcTotalHeating(...)    71.7508 kW

ratio between the two    1029.2

calcEnthalpy returns kJ/kg and the mass flow is kg/s, so their product is already kW. The ConSciR R source reads:

coolingPowerWatts <- massFlowRate * (h1 - h2)
coolingPowerKW    <- coolingPowerWatts / 1000

The name coolingPowerWatts gives the assumption away: that h is in J/kg. calcTotalHeating, in the same package, has no such division.

What to do: use abs(calcTotalHeating(...)) for a cooling load.

23.2 The sign in the VTT critical humidity

ConSciR and conscipy compute -0.00267·T³ + 0.160·T² + 3.13·T + 100, where Hukka & Viitanen (1999) have − 3.13·T:

actual output

 T (degC)   ConSciR / conscipy   Hukka & Viitanen
      5.0               119.32              88.02
     10.0               144.63              82.03
     15.0               173.94              80.04
     19.9               204.61              80.03
     20.0               205.24              80.04
     20.1                80.00              80.00
     25.0                80.00              80.00

A critical RH above 100 % can never be exceeded by a real reading,
so frac = (RHcrit - RH)/(RHcrit - 100) turns positive and M_max grows.
That is why a 12 degC / 53 %RH room accumulates a VTT index at all,
while the Zeng isopleth on the same readings reports no risk.

Look at the 20.0 and 20.1 rows. Above 20 °C the model uses a fixed 80 %; below it, the polynomial. The published coefficients make the polynomial land on 80.04 at 20 °C, joining continuously; +3.13T gives 205.24 before dropping to 80 — a 125-point discontinuity.

The consequence is concrete: when RH_crit exceeds 100 %, frac = (RH_crit − RH) / (RH_crit − 100) becomes a large positive number, M_max grows with it, and the model predicts growth in cold, dry air. Above 20 °C nothing is affected, and the Zeng isopleth is a separate implementation without the problem.

23.3 calcLM: three errors, two of which hide each other

The lifetime multiplier is meant to say how long something lasts in these conditions relative to 20 °C / 50 %RH. Michalski's rule of thumb is that 5 °C cooler roughly doubles the lifetime. What comes out instead:

actual output

Lifetime multiplier, as ConSciR and conscipy compute it:
   5.0 degC / 50 %RH  ->  0.997790
  10.0 degC / 50 %RH  ->  0.998552
  15.0 degC / 50 %RH  ->  0.999288
  20.0 degC / 50 %RH  ->  1.000000
  25.0 degC / 50 %RH  ->  1.000688
  30.0 degC / 50 %RH  ->  1.001354

A 5 degC drop should roughly double the lifetime (Michalski);
here 20 -> 15 degC changes it by -0.07 %.

                                      25C/50%   15C/50%   20C/30%
as written      EA=100 J   -  1/3      1.0007    0.9993    1.1856
units fixed     EA=100 kJ  -  1/3      1.9899    0.4907    1.1856
units and sign  EA=100 kJ  +  1/3      0.5025    2.0380    1.1856
published form  EA=100 kJ  +  1.3      0.5025    2.0380    1.9427

Cooling by 5 °C changes the lifetime by −0.07 %, and in the wrong direction — warmer reads as longer-lived. Three errors compound:

  • Units. EA=100 is treated as J/mol where the literature uses 100 kJ/mol, a factor of 1000. The exponential collapses to 1.000x and temperature all but drops out.
  • Sign. The code has exp(−EA/R · (1/T − 1/T₀)); the literature has +.
  • Humidity exponent. The code has (50/RH)^(1/3); the literature has ^1.3.
Fixing only the units makes it confidently more wrong

See the second row of that table. Changing EA=100 to 100e3 — the obvious fix — gives 1.99 at 25 °C: warmer, and twice the lifetime. The unit error was holding the sign error down at the 0.0007 level where nobody could see it; correcting the units alone amplifies it a thousandfold. All three have to change together. They then give 2.04 at 15 °C (5 °C cooler doubles the lifetime ✓) and 1.94 at 20 °C / 30 %RH (halving the humidity more than doubles it ✓) — both literature anchors land.

Until this is corrected upstream, do not read calcLM as a usable estimate of expected lifetime. calcPI is a separate formula — IPI's Arrhenius rate — and is unaffected.

23.4 The unit conversion in calcSH

actual output

calcMR(20.0, 50.0)              7.260814 g/kg dry air
calcSH(20.0, 50.0)              0.878947   <- MR / (1 + MR)

q = w / (1 + w) holds for w in kg/kg. With w in g/kg:
  w  = 7.260814 g/kg = 0.007260814 kg/kg
  q  = 0.007208475 kg/kg = 7.208475 g/kg

So the label says g/kg but the value is 8.2x too small.
The ordering check SH < MR still passes: True  -- which is why it never caught this.

The relation between specific humidity and mixing ratio, q = w / (1 + w), holds for w in kg/kg. But calcMR returns g/kg, so substituting directly loses the factor of 1000. The result is labelled g/kg and is 8.2 times too small.

The reach is wider than the library: add_humidity_calcs() writes an SH column every time, Orange's Humidity Calculations widget offers it as a tick box, and the Psychrometric Chart will put it on the vertical axis — a curve labelled g/kg and eight times too low.

This is exactly the limit of section 22

One of the internal consistency checks is SH < MR, and it has always passed — and it always will, because 0.879 genuinely is less than 7.26. An assertion about ordering cannot catch an error of scale. The check stays in this tutorial rather than being quietly removed, because nothing explains what verification does and does not buy you more clearly.

23.5 A missing temperature scores as “no risk”

actual output

A reading with no temperature, at 90 %RH:
  calcMould_Zeng(nan, 90)             nan
  calcMould_Zeng(nan, 90, label=True) 0.0   <- reads as 'no growth'
  calcMould_Zeng(20, nan, label=True) nan

mydata has 582 rows with no reading. tidy_TRHdata drops those,
but a logger that reports humidity while its thermistor is faulty
would not be dropped, and would be scored as safe.

calcMould_Zeng(..., label=True) checks whether the humidity is NaN but not the temperature. With no temperature every comparison is False and the result stays at 0.0 — and 0.0 means “no growth”. The DataFrame and the Orange widget then encode that as No risk, indistinguishable from a genuinely safe reading.

It should return NaN for the numeric result and a missing category, not No risk. tidy_TRHdata() drops rows where both readings are absent, so this only bites when humidity is fine and the thermistor has failed — which is precisely the case that most needs flagging.

How to weigh these five

conscipy is positioned as a faithful port of ConSciR, so reproducing upstream behaviour is deliberate: correcting it would break parity between the two. conscipy/tests/test_upstream_quirks.py pins the current behaviour so the documentation and the code fail together. But it also means this: the package suits teaching and exploration, and does not yet suit making real collection-risk or HVAC decisions — not from calcLM, calcSH, or the VTT index below 20 °C.

24References

If you use this in published work, please cite the original package and the framework it implements:

Shah, B., Cosaert, A., Beltran, V., et al. ConSciR: Tools for Conservation
Science. R package. https://bhavshah01.github.io/ConSciR/