Software development process¶
Development environment¶
We use Ubuntu 22.04 inside a Docker container as our development environment. Micromamba is used as a package management system and an environment is created using environment.yaml inside Docker.
Please see more details on developing inside a container using Visual Studio Code
Docker¶
Dockerfile overview¶
base: Ubuntu 22.04 base image, installs system packages, blobfuse2 and micromamba, installs the conda environment from environment.yaml and exposes the project at /workplace/SkiNet.
cpu:
FROM base AS cpu— installs CPU PyTorch wheels into the micromamba environment.gpu:
FROM base AS gpu— installs full (CUDA) PyTorch wheels.
Important notes
micromamba is used to create the
skinetconda environment. PATH is adjusted so the environment python is available.blobfuse2 is installed in the base image for convenience, but mounting Azure Blob Storage is usually done on the VM host and bind-mounted into containers.
If you need to run inside the container with FUSE support, you must run the container with appropriate capabilities and devices (see examples below)
images are labelled with the environment’s hash
Ways to build and run the containers¶
*** Build and push CPU and GPU images to the Hub ***
bash main_docker_build.sh
*** Manually build and push images to the Hub ***
To build and push an image to the Hub, run the following:
ENV_HASH=$(sha256sum environment.yaml | cut -c1-64)
IMAGE_TAG="pkliui/skinet:vtest" # adjust tag to match the version in on_start_gpu.sh / on_start_cpu.sh
docker build --no-cache --build-arg ENV_HASH=$ENV_HASH --target gpu -t $IMAGE_TAG .
docker push $IMAGE_TAG
*** Quick build & run commands ***
For quick experimenting and debugging purposes, use the following commands:
Build CPU image locally:
docker build --target cpu -t skinet:cpu .
Build GPU image locally:
docker build --target gpu -t skinet:gpu .
Run container bind-mount repo locally (and a host data dir at /mnt/data):
docker run -it \
-p 5000:5000 \
--mount type=bind,src=/Users/Pavel/Documents/repos/SkiNet,dst=/workplace/SkiNet \
--mount type=bind,src=/path/to/host/isic2017,dst=/mnt/data \
skinet:gpu bash
WARNING: When you launch a container manually like this, always bind-mount a host directory at /mnt/data before downloading data into it. If you run kaggle datasets download -p /mnt/data inside a container where /mnt/data is not mounted, the data is written to the container’s ephemeral writable layer and is lost the moment the container is removed (docker rm, or docker run --rm). Similarly, publish MLflow’s port with -p 5000:5000, or the UI started by start_mlflow.sh runs inside the container but is unreachable from the host. The startup scripts (on_start_gpu.sh / on_start_cpu.sh) handle both the mount and the port publish for you.
Additional options¶
To enable Azure Blob Fuse mounts, add
--cap-add=SYS_ADMIN --device=/dev/fuse --security-opt apparmor:unconfined
If you do not need FUSE inside the container (recommended): mount blobfuse on the VM host and only bind the mounted directory into the container; then you can omit the SYS_ADMIN/device flags.
Lightning Studio¶
Set up the environment on Studio¶
The following startup scripts are provided (each clones/updates the repo, pulls the Docker image, and launches a container):
Script |
Image |
Use case |
|---|---|---|
|
|
GPU training on a Lightning Studio GPU instance |
|
|
CPU development / debugging |
NOTE: When running code on Lightning Studio, data must be placed in Lightning Storage. The startup scripts bind-mount it into the container at /mnt/data/. Set azure_data: False and local_data_root: "/mnt/data/" in main_config.yaml. The host data directory defaults to Lightning Storage but is overridable per dataset — set ISIC_OUT_DIR (ISIC 2017) or PH2_DATA_DIR (PH2) to read data from any host path.
CPU modes (on_start_cpu.sh)¶
MODE |
What runs |
|---|---|
|
Drops into a |
|
|
The scripts bootstrap Docker and dispatch to the Python entry points. For what those entry points do — config fields, sweep mechanics, callbacks, reproducibility — see training.md.
GPU modes (on_start_gpu.sh)¶
RUN_TRAININGdefaults tofalse(dry-run guard). SetRUN_TRAINING=trueto actually pull the image and launch the container. Omitting it will print the resolved command and exit without running anything.
MODE |
What runs |
|---|---|
|
Drops into a |
|
|
|
|
|
|
|
|
|
|
Run CPU on Studio¶
# Interactive shell — ph2 dataset
DATASET=ph2 MODE=interactive bash on_start_cpu.sh
# Interactive shell — isic2017 dataset (default)
MODE=interactive bash on_start_cpu.sh
# Test a checkpoint — isic2017
MODE=test CHECKPOINT=runs/exp1/best.ckpt bash on_start_cpu.sh
# Test a checkpoint — ph2
DATASET=ph2 MODE=test CHECKPOINT=runs/exp1/best.ckpt bash on_start_cpu.sh
Run GPU on Studio¶
# Single training run — isic2017 dataset (default)
RUN_TRAINING=true MODE=train bash on_start_gpu.sh
# Single training run — ph2 dataset
RUN_TRAINING=true DATASET=ph2 MODE=train bash on_start_gpu.sh
# Multi-seed run using seeds/encoder/merge modes from YAML config
RUN_TRAINING=true MODE=seeds SEEDS="1 2 3 4 5" bash on_start_gpu.sh
# Multi-seed run with explicit encoder and merge modes (overrides YAML)
RUN_TRAINING=true DATASET=isic2017 ENCODER_MODES="classical" MERGE_MODES="classical local_refinement he2 attention_gate" MODE=seeds SEEDS="100 101 102 103 104" bash on_start_gpu.sh
# Optuna sweep — monitor and direction read from SWEEP_CONFIG in the YAML (recommended)
RUN_TRAINING=true MODE=sweep bash on_start_gpu.sh
# Optuna sweep — one-off override without editing the YAML
RUN_TRAINING=true MODE=sweep SWEEP_MONITOR=val_best_dice_at_threshold SWEEP_DIRECTION=maximize bash on_start_gpu.sh
# Threshold calibration
RUN_TRAINING=true MODE=calibrate bash on_start_gpu.sh
# Dry run — build command but skip container launch
RUN_TRAINING=false MODE=train bash on_start_gpu.sh
# Release GPU after training (default: NOT released)
RELEASE_GPU=true RUN_TRAINING=true MODE=train bash on_start_gpu.sh
After the container starts, attach to it in VSCode via the “Containers” tab → “Attach Shell”.
Debugging Cheatsheet¶
See ports busy with non-docker processes
Install lsof
sudo apt-get update
sudo apt-get install lsof
Example: port 6006
lsof -iTCP:6006 -sTCP:LISTEN -n -P
Kill the process
kill <PID>
Force save file from command line¶
sudo tee file_name.py > /dev/null << ‘EOF’ FILE-CONTENT EOF
Login to Codex¶
Login wih ChatGPT credentials
Authenticate wih ChatGPT as required, it will open a new window in your local browser. Note the port number
Top right corner SSH, click on “Connect via SSH” and it will issue you with a connection string
Modify it by adding relevant ports as follows for e.g. port 1455: ssh -N -L 1455:localhost:1455
@ssh.lightning.ai
Running tests¶
Tests use pytest. Run from the repo root:
# Run all tests
python -m pytest Tests/
# Run a specific module
python -m pytest Tests/ML/configs/
# Run with verbose output
python -m pytest Tests/ -v
# Run with coverage
python -m pytest Tests/ --cov=SkiNet --cov-report=term-missing
# Skip slow forward/backward passes on large inputs (the `slow` marker, defined in pytest.ini)
python -m pytest Tests/ -m "not slow"
pytest.ini sets testpaths = Tests and addopts = -q --tb=short, and registers the slow and
unet2d markers.
The test suite is organized to mirror the source tree:
Tests/ML/configs/– Pydantic config validationTests/ML/datasets/(andTests/ML/datasets/preprocessing/) – datasets and CSV buildersTests/ML/dataloaders/– DataLoader behaviourTests/ML/model/(witharchitecture/andblocks/) – UNet2D architecture and blocksTests/ML/transformations/– augmentation pipelinesTests/ML/training/– loss and training utilitiesTests/ML/utils/– config loading, samplingTests/Utils/– analysis, data, logging, MLops helpersTests/Azure/– Azure integration (requires credentials)
conftest.py at the repo root defines shared fixtures.
Linting and type checking¶
# Flake8 (configured in .flake8)
flake8 SkiNet/ Tests/
# mypy via the repo wrapper (uses mypy.ini; defaults to the SkiNet and Tests packages)
python check_types.py
# run every hook against the whole tree
pre-commit run --all-files
Pre-commit hooks are defined in .pre-commit-config.yaml (all language: system, i.e. run in your
active env). Three hooks run:
Hook id |
Command |
Notes |
|---|---|---|
|
|
Python files only |
|
|
mypy via the wrapper; |
|
|
|
Install once with:
pre-commit install
Kaggle notebooks¶
on_start_kaggle.sh runs SkiNet directly in a Kaggle GPU session — no Docker. It
clones/updates the repo, starts MLflow, generates the dataset metadata CSV, and dispatches to a
Python entry point. The ISIC 2017 data is expected to be attached to the notebook as a Kaggle
dataset (mounted under /kaggle/input/...).
Variable |
Values / default |
Notes |
|---|---|---|
|
|
Maps to |
|
|
Selects data dir and config |
|
space-separated, required for |
|
|
space-separated; omit to use YAML |
Override architecture sweep modes |
|
unset → dry run |
The script executes only when |
# Single ISIC 2017 training run on Kaggle
RUN_TRAINING=true MODE=train bash on_start_kaggle.sh
# Multi-seed run with explicit architecture modes
RUN_TRAINING=true DATASET=isic2017 ENCODER_MODES="classical" MERGE_MODES="he2 attention_gate" \
MODE=seeds SEEDS="100 101 102" bash on_start_kaggle.sh