A multicore chip has no temperature sensors, so the controller has to model each core's temperature from its load and the cooling it gets, shut cores down before they cook, and decide which cores get active cooling. Build OverheatController.
API
OverheatController(passive_capacity, active_per_core, core_ids): 1 to 1,023 cores with unique string ids. Every core starts at 20.0 °C, running, load 0, no active cooling, statusidle. Both capacities are positive (watts).set_core_load(timestamp, core_id, load): record a pending load (watts,0 <= load <= 32768). Nothing happens until the nexttick. If a core gets several calls between ticks, only the last one counts. On a shut-down core, the call is a restart request (see step 3).tick(timestamp) -> list[str]: run the five steps below and return the status changes.get_temperature(core_id) -> float: the core's temperature as of the last tick.
Timestamps are integer milliseconds and strictly increase across all calls. Inputs are always valid.
Cooling model
- Passive cooling is shared. Each core (running or not) asks for
load + 2watts. Withkcores under active cooling, the shared capacity isC = passive_capacity * (1 - penalty(k) / 100), wherepenalty(k) = 10/1 + 10/2 + 10/3 + 10/5 + 10/8 + ...summed over the firstkFibonacci terms1, 2, 3, 5, 8, 13, ...(penalty(0) = 0). If total demand is at mostC, every core gets its full demand; otherwise each core getsC * its_demand / total_demand. - Active cooling gives each selected core another
active_per_corewatts. - Over an interval of
dtseconds a core's temperature changes by0.02 * net * dt, wherenet = load - passive_share - active_share, and it never drops below 20.0.
tick(timestamp) in order
- Advance every core from the previous tick to
timestampin one step, using the loads and active-cooling selection that were in force during that interval. (On the first tick there's nothing to advance.) - Shut down every running core whose temperature is now
>= 80: it stops running and its load becomes 0. - Apply pending loads. A running core takes its new load. A shut-down core with a pending load restarts with that load only if its temperature is strictly below 50; otherwise the request is dropped (it stays shut down, and it takes a new
set_core_loadto try again). Clear all pending loads. - Choose active cooling for the next interval. Start with no core selected. A running, unselected core must be selected if its temperature is
> 60, or if its rise rate0.02 * (load - passive_share)(with passive shares computed for the current number of selected cores, and no active cooling for itself) is> 0.5°C/s. Selecting cores raises the penalty, which shrinks everyone's passive share and can push more cores over the rate threshold, so repeat until a full pass selects nobody new. One pass is not enough. - Report. A core's status is
shutdownif it isn't running,coolingif it was just selected, andidleotherwise. Return"core_id=status"for every core whose status differs from its status after the previous tick (initiallyidle), sorted bycore_idas strings. Return[]if nothing changed.
The tests stay far from every threshold, so float rounding won't flip a comparison. Temperatures are checked to within 1e-6.
With 1,023 cores and hundreds of ticks, each pass of step 4 must be linear: compute the total demand once per pass, not once per core.
c = OverheatController(10, 5, ["c0"])
c.set_core_load(1000, "c0", 60)
c.tick(1000) # ["c0=cooling"]: demand 62 > 10, share 10, rate 0.02 * 50 = 1.0
c.tick(3000) # []: 2 s at net 60 - 9 - 5 = 46 (penalty(1) = 10%, so C = 9)
c.get_temperature("c0") # 21.84
c.tick(71000) # ["c0=shutdown"]: 21.84 + 0.92 * 68 = 84.4
c.set_core_load(72000, "c0", 10)
c.tick(72001) # []: about 84.36, not below 50, so the restart request is dropped
c.tick(1000000) # []: cooled to about 47.24, but no request is pending
c.set_core_load(1000001, "c0", 10)
c.tick(1000002) # ["c0=idle"]: restarted with load 10