Skip to content

Instances, ports and leases

ArduPilot’s SITL takes an instance number, -I n, and derives every port from it. The MAVLink TCP port is 5760 + 10n, with SITL’s further serial ports just above it, and the JSON physics backend is 9002 + 10n. Instance 0 is 5760, instance 40 is 6160.

So two SITLs given the same number on one machine collide. At best the second fails to bind. At worse a client connects to a port that a different flight’s SITL already holds, and flies someone else’s aircraft. On a workstation where several agents, a test suite and a person at a terminal all start SITLs, picking a number is not a private decision.

mcardupilot gives each instance in the pool a lock file, leases/instance-<n>.lock under the data directory. Holding an exclusive flock on that file is holding the lease. The holder writes its name, pid and the time into the file, so anyone can ask who has instance 44 and get an answer like mcardupilot session f391b1 (tools-hop-test) pid 728436.

The property that decides it is what happens on death. The kernel drops a flock when the process holding it exits, crash and SIGKILL included. There is no expiry to tune, no heartbeat to refresh, no stale-lock sweep to run, and no way for a dead holder to strand an instance. A lock file with a pid in it but no lock on it is simply free.

Two smaller properties help. A flock belongs to the open file description, so two leases taken inside one process exclude each other just as leases in different processes do; a test suite running several SITLs in one interpreter cannot double-book. And the lock is advisory, so reading the file to name the holder needs no cooperation from it.

A lock only excludes callers that take locks. autotest.py, which always uses instance 0, and make fly take none. So every launch probes the instance’s MAVLink and JSON ports while holding the lock, and skips an instance whose ports are busy, naming the process that holds them. Instance 0 is refused in any pool for the same reason: its owner never asks.

Instances 14 and 15 are refused as well. Their MAVLink ports, 5900 and 5910, land in the range VNC servers use, which is a common sight on a machine that also serves remote desktops.

Every agent session runs its own mcardupilot server, and all of them share one data directory. Each session’s SITL writes sessions/<n>/sitl.pid with its process group and start time.

When a server starts, it wants to kill SITLs that a crashed server left behind. A pid file is not evidence of that: the directory is full of pid files belonging to other servers’ live sessions. So for each pid file it tries to take that instance’s lease. If it can, nothing alive holds the instance, and whatever the pid file names is an orphan. If it cannot, a live server or script owns the instance, and it is left alone.

The same test guards kill_instance. A SITL on an instance someone holds is refused with the holder named; only an instance nobody holds can have its stray killed, and the killer holds the lease for the whole kill so nothing new starts there in between.

A pid by itself can also be recycled. The pid file records the leader’s start time in clock ticks from /proc/<pid>/stat, and a process whose start time differs is someone else’s and is left alone.

The pool, 40-59 by default, is only which numbers a server or script will try. It is one setting, so a harness can widen its own pool without touching anyone else’s. What must be shared is the lock directory: list_instances and python -m mcardupilot.leases status read every lock file in it, inside the pool or not, because a held lease nobody can see is exactly what leases exist to prevent.