Reading Settings
Font Size
16px
Line Spacing
1.6
Reading Width
900px
Font Share
Theme
Text To Speech

178: Chapter 178 Keeping a Calorie Account

The twelfth trip to the Wasteland, July 3rd of the 24th year.

Outpost Workroom Number Two.

Four servers were lined up along the wall, and twenty-seven 170HX boards were distributed across four groups of nodes, their fans kept at minimum speed, emitting a patient hum.

They had already learned how to deceive.

Without taking them apart, no one would have guessed that this batch of boards, which should have been quietly staying in the mining farm calculating hash values, had been pried open all the way from their firmware, drivers, and runtime by Jiang Lin, piecing together nearly one TiB of stable, computable HBM.

The work assigned to them by the mining farm was very simple: moving the same kind of bricks day after day.

In the hands of Jiang Lin, matrix block division, model inference, code indexing, and proof dependency graphs could all be stuffed inside.

However, between a batch of boards being able to compute and a computing cluster being able to work long-term, there lay a very hot threshold.

Jiang Lin pulled the file created four days ago onto the main screen.

[MPS-ThermalFabric_v0.1]

The local service package handed over by TM-7 had already been transferred to read-only archival storage.

Future device IDs, medium parameters, and original service protocols were all isolated outside the experimental chain, and the control system could only read the interfaces redefined by Jiang Lin.

The original thermal management table of the maintenance port was very simple.

Each branch line was followed by a number in megajoules, representing how much more heat it could swallow.

This style of writing was suitable for phase-change buffer units.

That thing was like a pre-emptively emptied reservoir; how much storage capacity remained and how large of a flood peak it could handle were clear at a glance.

The liquid cooling loop cobbled together with water and ethylene glycol in front of him was much more troublesome; while it absorbed heat, it also sent the heat away to the terminal heat exchanger.

Recording only the remaining capacity was like only staring at how much empty space was left in a warehouse while completely forgetting how much cargo went in and out every day.

Thus, Jiang Lin opened two ledgers for each branch line.

[Continuous Heat Rejection Capacity: kW]

[Transient Thermal Budget Within Response Window: MJ]

The first ledger managed the long term.

How much heat a branch line could carry away per second determined whether a task could run for an hour or for a year.

The second ledger managed the immediate present.

It took time for the pump group to speed up, time for the valves to turn, and likewise time for the coolant to travel from the main pipe to the cold plate.

During these dozens of seconds, who caught the heat peak already generated by the chip determined whether the cards would continue computing or throttle themselves first due to overheating.

Computing power had queues, video memory had windows, and faults had ledgers.

Now, heat had to be entered into the ledger as well.

Jiang Lin cut off the last connection between the local service package and the experimental control network, and began modifying the four servers.

The original air ducts were retained to continue taking care of the motherboards, power supplies, storage, and other components without pressed cold plates.

The GPU and HBM areas of the twenty-seven 170HX cards were covered with copper cold plates respectively, and the four groups of nodes corresponded to four parallel liquid cooling branch lines.

Three circulation pumps were responsible for the common water supply and return, and a normally closed redundant bypass was added next to each main line.

Under normal conditions, the bypass was closed.

Only after the flow of the main line was cut off did it qualify to open.

At the end of the branch lines was a set of finned liquid-to-air heat exchangers, and twelve industrial fans were responsible for blowing the hot air out of the workroom.

Inlet temperature, outlet temperature, flow rate, and supply and return water pressure all had independent measurement points.

The highest temperature among the temperature measurement results of the GPU core, HBM area, and on-board power supply area was taken and uniformly recorded as the card hotspot temperature.

After two abandoned pressure vessels completed cleaning and initial flaw detection, they were dragged into the workroom.

Jiang Lin welded metal fins inside, completed the packaging, re-conducted weld flaw detection, hydraulic and airtightness tests, and finally poured in a water-ethylene glycol mixture, converting them into sensible heat buffer tanks.

These two tanks were obviously far from advanced.

They were big and heavy, and their heat absorption capacity per unit volume was also very ordinary. Placed alongside the phase-change buffer unit of TM-7, it was roughly equivalent to an iron bucket participating in a future industrial product award ceremony.

But iron buckets had the advantages of iron buckets.

Materials could be bought, welds could be inspected, broken ones could be disassembled, and once disassembled, another could be made just like it.

The Real World could handle such things.

After the pipelines were completed, Jiang Lin still did not let the thermal fabric take over control.

He first left the twenty-seven cards in standby mode, connected adjustable resistance heating blocks to the four sets of cold plate loops, and poured heat into the coolant gear by gear.

Six kilowatts.

Eight kilowatts.

Ten kilowatts.

Eleven kilowatts.

Eleven point eight kilowatts.

At the final gear, the liquid outlet temperature of the terminal heat exchanger slowly crept upward.

Two hours later, the curve stopped and climbed no higher.

Jiang Lin noted the results on the test page.

[Terminal Continuous Heat Rejection Capacity: 11.8 kW (at current incoming air temperature and rated air volume)]

This was not a constant that stood independent of the environment.

The incoming air temperature, fan speed, and liquid-side flow rate were written into the calibration records together.

[Four-Node Cluster Estimated Full-Load IT Power: 9.6 kW to 10.1 kW]

The heat exchanger was sufficient, and the total flow rate of the pump group was also sufficient.

This was good news, even suspiciously good.

Even though the hardware clearly could carry away the heat of the entire cluster, the four servers had previously consistently failed to run at full capacity simultaneously.

Where did the heat go?

July 17th of the 24th year, the fixed valve position baseline test began.

The four branch lines completed static hydraulic balancing according to the rated thermal load.

The valve openings were written into the records and completely locked during the test.

Twenty-seven 170HX cards established shadow windows, and a unified task queue entered the four groups of nodes sequentially.

For the first twenty minutes, the curves were very pretty.

At the thirty-first minute, the card hotspot temperature of NODE-D first crossed 75 degrees Celsius.

This group of nodes undertook more checkpoint compression and video memory transfer; the GPU core still had leeway, but the HBM and on-board power supply areas had already begun to accumulate heat.

The return water temperature of the fourth branch line edged upward, while the second and third branch lines remained cool, and NODE-B even had nearly half of its cooling margin sitting idle.

The 11.8 kW terminal capacity was still there.

It was just evenly divided among the four pipes, and no one ever asked which group of nodes generated more heat today.

At the forty-sixth minute, the highest card hotspot temperature of NODE-D reached 79.8 degrees Celsius, and the first card triggered frequency throttling.

Jiang Lin continued to wait.

Lowering NODE-D alone was useless.

This batch of matrix blocks had a synchronization barrier; if it slowed down, the other three groups of nodes had to stand and wait with it in front of the checkpoint.

Migrating the active video memory window at this time would also require re-establishing mappings.

MPS-Scheduler could ultimately only chop down the concurrency of the entire batch of tasks gear by gear.

Ninety percent.

Eighty percent.

Seventy percent.

When the total load dropped to sixty-four percent, the temperature curve of NODE-D was finally flattened.

The remaining three branch lines still had margins.

The fixed valve position mode continued to run for six hours; the equipment remained intact, and over-temperature protection was not triggered even once.

The cost was equally clear: normalized against a full-load IT power of 9.8 kW, under fixed valve positions, the four nodes could only release 64% of their safe continuous computational load.

[Fixed Valve Position Baseline]

[Test Hardware: 27 170HX Cards]

[Terminal Heat Rejection Capacity: 11.8 kW (same as calibration working condition)]

[Four-Node Cluster Full-Load IT Power Baseline: 9.8 kW]

[Safe Continuous Load Upper Limit: 64%]

[Source of Restriction: Local Branch Heat Accumulation]

[Idle Cooling Capacity: Exists]

The four branch lines clearly added up to be sufficient, yet the hottest one still could not consume the remaining cooling capacity from elsewhere.

There was plenty of food in a warehouse, but the starving people stood behind another locked door.

It was meaningless to keep piling grain into the warehouse; the door had to be opened first.

Jiang Lin wrote down the problem definition on the front page of the baseline report.

[Total cooling system capacity is sufficient; local cooling capacity cannot transfer with computational tasks.]

There were two easiest remedies.

Re-adjust the static hydraulic balance, or increase the total flow rate and fan speed according to the demands of the hottest node.

The former scheme only recognized one stable load; whenever the task changed, the balance had to be redone accordingly.

The latter scheme was more straightforward; as long as one was willing to pay for power consumption, pump margins, and heat exchange redundancy in the long run, the worst-case operating conditions could always be suppressed.

Engineering frequently did this; it was reliable, saved brainpower, and its main drawback was that it cost money.

Mature computing power centers were long equipped with variable-frequency pumps, temperature control valves, and thermal-aware scheduling.

What Jiang Lin was pushing forward right now was not merely adding an automatic switch to the water supply pump.

He wanted to compress the heat peaks about to be generated by tasks, the in-transit heat not yet delivered by the pipelines, branch recovery time, and checkpoint migration qualifications all into the same set of verifiable operational states.

Task planning, in-transit heat, valve response, and checkpoint migration had to enter the same state table.

Cooling capacity needed to be measurable, reservable, and migratable just like computational tasks, and it also needed to know when to refuse.

August 12th of the 24th year, the MPS thermal fabric took over the four branch lines for the first time.

NODE-A and NODE-B entered the computing state, while the other two groups remained on standby.

This round only verified joint flow regulation under stable operating conditions: after the two groups of nodes entered a stable load, the card hotspot temperatures had to be maintained below 75 degrees Celsius.

Four minutes and twenty seconds after the task started, the temperature rise slope of NODE-A exceeded predictions.

The valve of the first branch line opened wider.

The flow rate increased, the branch pressure rose, the pressure difference of the common trunk pipe changed accordingly, and the flow rate obtained by NODE-B began to decrease.

Forty seconds later, the temperature of NODE-B ticked upward.

The system turned to open the valve of the second branch line.

NODE-B was saved, and NODE-A started heating up again.

The two branch lines started a tug-of-war around the same set of pumps.

The solenoid valves bounced back and forth between 34% and 71% openings; the pressure curve moved first, followed by the flow curve, with the temperature curve dragging at the very end. By the time the temperature finally delivered the message to the controller, the previous round of commands had long since changed the situation beyond recognition.

At the eleventh minute, the first 170HX of NODE-B throttled its frequency.

At the thirteenth minute, the second card approached the protection boundary.

Jiang Lin pressed the abort button.

The fans were still spinning, and the valves were still busy.

They strictly executed every single command; the problem lay precisely in the commands being too eager.

[First Joint Flow Regulation Test]

[Result: Failed]

[Fault Type: Branch Coupled Oscillation]

[Equipment Loss: 0]

The original logic of TM-7 faced high-speed valve groups, precise pressure difference control, and phase-change buffering.

The branch lines there had just received instructions, and new cooling capacity was already on the way.

Real solenoid valves were more observant of rules.

It took several seconds from receiving a command to stabilizing the flow rate, and after the circulation pump changed its speed, the entire loop had to find its balance all over again.

Stuffing the control rhythm of a future system into it as-is was like using fighter jet motion commands to urge a fully loaded forklift.

The forklift was already trying very hard, but the cargo would still spill.

Jiang Lin added three constraints to the controller.

[Any flow rate adjustment must wait for the current branch line to complete its response.]

[Prohibit continuous reverse correction based on a single temperature value.]

[Pressure changes must enter the predictive model prior to temperature changes.]

December of the 24th year, the second round of control logic went online.

Each of the four branch lines now had its own dynamic ledger.

Continuous heat rejection capacity, response window thermal budget, coolant quality, inlet and outlet temperatures, flow rate, cold plate thermal lag, and buffer tank status were written in item by item.

The first forty minutes of the second test were much smoother.

NODE-A and NODE-B remained stable. After NODE-C started, the third branch line still had a 38% margin on paper.

Based on this, the system approved the next group of matrix tasks to enter.

Seven minutes later, three consecutive cards in NODE-C crossed the warning temperature.

Outlet temperature was normal.

Flow rate was normal.

The margin on the ledger was still there.

Jiang Lin stared at the curves for a few seconds and shut down the task.

That 38% margin had never existed from beginning to end.

The heat just generated by the chip first enters the thermal conductive layer, then enters the copper cold plate, hoses, and coolant, and finally reaches the branch outlet.

The temperature seen by the sensor belongs to more than a hundred seconds ago, and the ledger approved new tasks with old news.

Simply put, the heat has already set off, but the report still treats it as if it hasn't hit the road.

[Second Joint Scheduling Test]

[Result: Failed]

[Fault Type: In-Transit Heat Not Counted]

[Frequency Reduction Triggered: 3 Cards]

[Equipment Loss: 0]

Jiang Lin removed the cold plate of NODE-C and laid out measurement points all along the entire thermal conduction path.

Near the chip.

Cold plate inlet.

Cold plate outlet.

Branch return water.

Buffer tank inlet.

Terminal heat exchanger.

The same amount of heat left six timestamps at six locations.

From the appearance of the chip temperature rise to the return water temperature being sufficient to stably reflect the load, there is a difference of one hundred and seventeen seconds.

One hundred and seventeen seconds is enough for the switch to move many batches of data, and enough for three 170HX cards to force themselves into the frequency reduction zone.

For water, this is just the normal journey of a pipeline.

The response window thermal budget was rewritten accordingly.

[Response Window Thermal Budget = Available Sensible Heat Capacity Allocated to This Branch + Expected Heat Dissipation within the Window - In-Transit Heat - Safety Margin]

Two buffer tanks are shared by four branches, and the same capacity cannot be recorded in four ledgers at the same time. The system must first reserve capacity for each branch, and then update the available portion based on the tank temperature, flow rate, and recovery status.

Allocated capacity and expected heat dissipation within the window can be estimated through calibration and sensors; the safety margin is preset.

The most difficult part is the in-transit heat.

It hides in the thermal conductive layer, cold plate, hoses, and flowing liquid. It cannot be directly read by a single measurement point, nor will it disappear on its own just because the report temporarily cannot see it.

Jiang Lin sent the power consumption curve, window mapping scale, kernel type, and estimated duration of each 170HX into MPS-Scheduler.

Matrix calculations have their own thermal peaks.

Model inference has its own thermal peaks.

The average power consumption of log indexing and checkpoint writing is similar, but their short-term pulses are completely different.

Starting from this round of iteration, the computing power scheduler will submit task plans in advance.

Before the task even enters the GPU, the heat it will generate in the next few minutes is already recorded in the branch ledger ahead of time.

In April of Year 25, the third round of control logic entered testing.

Thirty seconds before the matrix task of NODE-A started, the first branch began to increase flow.

When the task truly entered calculation, the coolant had already arrived at the cold plate.

The second group of tasks was originally prepared to be sent to NODE-C.

The surface temperature of NODE-C was lower at this moment, but the ledger showed that its buffer recovery speed was slower than that of NODE-B.

The system abandoned this seemingly cooler node and sent the task into NODE-B, which had a slightly higher temperature but a more sufficient thermal budget.

Twenty minutes later, neither branch crossed the line.

Phase one passed.

Jiang Lin immediately closed 70% of the main return water valve of NODE-B, simulating a main return water path failure.

The flow meter reported an error first.

The ThermalFabric froze new tasks for NODE-B, the running chunks began to write checkpoints, and NODE-A and NODE-C received migration plans.

The order is correct.

Seventeen seconds later, the first card of NODE-A still experienced frequency reduction.

Crossing the switch took less than a second for the task.

Before the original chunk of NODE-A ended, a new shadow window had already been established, and the computing power load rose instantly.

From receiving the plan to establishing effective flow, the corresponding branch required twenty-six seconds.

At the ninth second, the temperature rise slope crossed the line.

At the seventeenth second, frequency reduction was triggered.

Jiang Lin terminated the migration.

[Third Fault Migration Test]

[Result: Failed]

[Fault Type: Computing Power Migration Faster Than Cooling Response]

[Lost Checkpoints: 0]

[Frequency Reduction Triggered: 1 Card]

Data can take shortcuts, but water must honestly walk through the pipes.

Before cooling capacity arrives, idle computing power is merely a paper asset.

Whoever crams the task in first will receive an over-temperature bill.

Jiang Lin added a door in front of the task migration state machine.

[COOLING_READY]

Only when all five conditions are met is the target node allowed to receive the migration load.

Branch flow reaches the planned value.

Inlet temperature enters the allowable range.

Supply and return water pressure difference enters a stable interval.

In-transit heat is lower than the safety boundary.

The buffer tank has reserved capacity for new tasks.

Computing nodes wait.

Pumps and valves go first.

In the autumn of Year 25, MPS-ThermalFabric completed its seventh internal iteration.

A set of thermal states was added behind each of the four groups of nodes.

[NODE-A]

[Available VRAM: 284 GiB]

[Available Calculation Window: 7]

[Response Window Thermal Budget: 31.2 MJ]

[Estimated Recovery Time: 18 min]

[Fault Migration In: Allowed]

[NODE-B]

[Available VRAM: 301 GiB]

[Available Calculation Window: 8]

[Response Window Thermal Budget: 8.7 MJ]

[Estimated Recovery Time: 43 min]

[Fault Migration In: Denied]

NODE-B had more VRAM and more idle calculation windows, but the task was still blocked outside the door.

For the first time, the thermal budget acquired the power to veto computing power scheduling.

This power was repeatedly used over the next half year.

High-temperature startup.

Low-temperature startup.

Pump group speed reduction.

Fan failure.

Filter clogging.

Flow sensor drift.

Single-branch leakage simulation.

Server fan standstill.

Checkpoint writing delay.

The filter cotton in Workshop No. 2 was replaced batch after batch, the valve actuators were disassembled seven times, and new measurement points and weld repair marks appeared on the outer walls of the two buffer tanks.

The Wasteland walked from autumn to winter, and from winter into the spring of Year 26.

Seventy-six sets of fault scenarios were fed into the system, leaving seventy-six records.

The MPS ThermalFabric never guarantees that every task can continue.

Insufficient evidence, no load increase.

If the thermal budget cannot cover the next stage, the task is frozen at the checkpoint.

If the branch recovery time is longer than the task window, the load is transferred to other nodes.

All safe paths are blocked, and the task exits the queue.

The capability of a reliable system includes both finishing the work and knowing which work cannot be done at the moment.

On May 19 of Year 26, the final control verification began.

Six days prior, the Outpost shut down large processing equipment, and the photovoltaic array continuously charged the energy storage.

Two Wasteland fans completed bearing and pitch inspections, and the 80 kWh energy storage reached a 93% state of charge.

Jiang Lin split the final test into two phases.

Twenty-seven 170HX cards.

Four servers.

Four liquid cooling branches.

Three circulation pumps.

Two thermal buffer tanks.

The same set of valves, sensors, cold plates, and terminal heat exchangers.

The task queue was also completely identical.

Phase one locked the valve positions.

Phase two connected to MPS-ThermalFabric.

No hardware was replaced during the test.

[Stable Operating Condition Board Hotspot Temperature Upper Limit: 75°C]

[Flow Interruption Fault Transient Board Hotspot Temperature Upper Limit: 76°C]

Even the weather had to be as identical as possible.

An adjustable mixing windbox was connected to the air inlet end of the terminal heat exchanger.

The eight-hour air inlet temperature curve experienced by the fixed valve position group would be completely recorded and reproduced in the same order two days later.

The initial water supply temperature, branch pressure, task initial checkpoint, and energy storage state of charge all had to return to the same window.

Otherwise, if one side encountered cool weather and the other encountered hot wind, no matter how beautiful the result, it would only be fit to fool oneself.

At 7:30 AM on May 19, the fixed valve position group started.

To keep NODE-D under control, the cluster load was limited to 64%.

Eight hours later, the task queue completion rate was 63.7%.

There was no board over-temperature, and the average capacity of the heat exchanger was only used by 67%.

[Fixed Valve Position Mode]

[Cluster Safe Continuous Load: 64%]

[Task Queue Completion Rate: 63.7%]

[Over-Temperature Protection Triggered: 0]

[Automatic Bypass Takeover and Task Linkage: Disabled]

[Terminal Heat Exchange Capacity Utilization Rate: 67%]

The test ended, the servers shut down, and the circulation pumps continued to run.

After waiting for the heat retained in the cold plates, hoses, and buffer tanks to completely drop back, the energy storage was replenished to 93%, and the task queue was restored to the same initial checkpoint.

At 8:00 AM two days later, the second phase began.

MPS-ThermalFabric took over the cluster.

The twenty-seven 170HX cards re-established shadow windows.

The same batch of matrix chunks, model inference, code indexing, and proof dependency graphs entered the queue.

[Stable Computable HBM: Approx. 1 TiB]

[Target Duration: 8 h]

[Allowed Over-Temperature Protection: 0]

[Allowed Lost Checkpoints: 0]

The base matrix blocks, model weights, and index files had already established read-only shadow mappings on the backup nodes.

MPS-Checkpoint only saves the incremental state of the current chunk.

When a fault occurs, the only thing that truly needs to be moved away is the changes since the last checkpoint, rather than the entire VRAM group of NODE-B.

In the first hour, the cluster IT load stabilized around 9.8 kW.

Each of the four branches ran its own flow rate.

NODE-A undertook continuous matrix calculations, with the highest cooling flow.

NODE-C focused on short-term model tasks, handing over thermal peaks to the buffer tank; NODE-D had the largest internal board differences, always keeping a thicker safety margin on the ledger.

In the second hour, the reproduced air inlet curve entered the warming-up phase.

The terminal heat exchange efficiency decreased.

The system delayed two groups of low-priority tasks by seventeen minutes, waiting for the first buffer tank to recover before putting them back into the queue.

The twelve fans were not brutally ramped up to maximum simultaneously.

In the third hour, the reading of a flow sensor inexplicably drifted downward by eight percentage points.

EvidenceGate checked the supply and return water pressure difference, pump speed, and valve position feedback; none of the three data points supported that the branch flow had actually decreased.

The drifted sensor was downgraded to a low-confidence source, and the controller continued to execute the original plan.

At 5 hours and 12 minutes, Jiang Lin walked to the side of the main circuit and gripped the shut-off valve of the main return water path of NODE-B.

In the previous seventy-six groups of fault tests, this action had been performed many times.

The final verification only recognizes this single time.

The valve dropped to the bottom.

The flow rate of the second branch quickly dropped to zero.

[BRANCH-B Main Path: Flow Lost]

[Suspected Fault: Waiting for Pressure/Flow Dual Evidence]

[Freeze New Task Dispatch]

[Initiate Current Checkpoint Writing]

The eight cards on NODE-B were still running. The residual coolant in the cold plates and branches gave them a short thermal inertia window; calculated at the current load, they wouldn't last two minutes.

At the sixth second, the current chunk ended.

The supply and return water pressure difference crossed the fault confirmation threshold simultaneously.

[Fault Confirmation: Pressure/Flow Dual Evidence]

[Start NODE-B Backup Bypass]

At the ninth second, checkpoint writing was completed.

At the eleventh second, NODE-A and NODE-D received the pre-migration plan.

Neither group of nodes immediately received the migration load.

The backup bypass of NODE-B began to establish return water capacity.

The circulation pumps increased speed synchronously, and the two migration target branches, NODE-A and NODE-D, increased flow in advance to establish new hydraulic states respectively.

[NODE-A: Waiting for COOLING_READY]

[NODE-D: Waiting for COOLING_READY]

At the eighteenth second, the inlet flow of NODE-D reached the planned value.

At the twenty-first second, the pressure difference of NODE-A stabilized.

At the twenty-third second, both groups of statuses turned green.

[COOLING_READY]

Migration began.

The matrix chunk with the longest remaining running time was transferred first, followed by model inference. Log indexing was frozen at the checkpoint, and proof dependency graph scanning was downgraded to the lowest priority.

The load of NODE-B dropped rapidly, while the board hotspot temperature continued to rise by inertia.

73 °C.

74.2 °C.

75.1 °C.

The steady-state temperature threshold had been crossed, leaving only 0.9 °C until the fault transient upper limit.

At the forty-sixth second, the backup bypass completed its takeover, and the new coolant entered the NODE-B cold plate.

The temperature rise curve stopped.

At the seventy-first second, the curve began to recede.

[Main return water pathway: Isolated]

[Backup bypass: Taken over]

[Task migration: Completed]

[Lost checkpoints: 0]

[Overtemperature protection triggered: 0]

[Cluster computing state: Continuing]

Jiang Lin did not reopen the shut-off valve.

For the next two hours and forty-eight minutes, NODE-B ran entirely on the backup bypass.

During the migration, the cluster's total computing power dropped briefly by 13 percent.

After the cooling state stabilized, some tasks returned to NODE-B.

At four in the afternoon, the eight hours came to an end.

The last matrix block was written to storage.

Twenty-seven 170HX cards exited the computing state one by one, and the circulation pump continued to run for thirty-seven minutes.

Computing tasks could be stopped on a whim, but heat did not accept administrative orders.

It remained in the cold plates, hoses, and buffer tanks, and had to be sent out section by section.

The power supply log was generated simultaneously.

[Average photovoltaic and fan output during testing: 8.3 kW]

[Average total electrical load of the test system: 10.5 kW]

[Energy storage purpose: To shoulder power gaps and switching fluctuations]

[State of charge of energy storage at test end: 70%]

The eighty-kilowatt-hour energy storage only filled the gap between power generation and load, while photovoltaics and wind turbines bore the majority of continuous power supply.

If the entire test were attributed to energy storage, the account would fall apart at a second glance.

The final report unfolded on the main screen.

[MPS-ThermalFabric_v0.1 Final Verification]

[Test hardware: Completely identical to fixed valve position mode]

[Newly added pump groups: 0]

[Newly added heat exchangers: 0]

[Newly added cold plates: 0]

[Fixed valve position mode safe continuous load: 64%]

[Thermal Fabric mode average effective load: 96.8%]

[Same task queue completion rate: 97.6%]

[Average effective load improvement relative to fixed valve position baseline: 51.2%]

[Simulated main return water pathway complete flow interruption: 1]

[Fault confirmation time: 6s]

[Checkpoint completion time: 9s]

[Cooling readiness and task migration initiation: 23s]

[Backup bypass takeover: 46s]

[Overtemperature protection triggered: 0]

[Lost checkpoints: 0]

[Unrecoverable tasks: 0]

[Flow interruption fault transient temperature upper limit: 76 ℃]

[Measured highest board hotspot temperature: 75.1 ℃]

[Result: Passed]

The twenty-seven cards remained unchanged.

The pump groups, valves, cold plates, and heat exchangers also remained unchanged.

Fixed valve position mode could only safely release sixty-four percent of the continuous load, while Thermal Fabric pushed the average effective load to ninety-six point eight percent.

From start to finish, the system did not generate a single extra watt of cooling capacity.

It only did three things.

Knowing in advance where heat would be generated, reserving a budget for that branch to catch the heat peak, and allowing the cooling capacity to reach the target node before the migration task.

Now, computing power, VRAM, and cooling capacity finally appeared on the same scheduling table.

Jiang Lin wrote the application boundaries at the end of the report.

[The above improvements only apply to the current four-node heterogeneous cluster.]

[They must not be directly extrapolated to other cabinets, boards, or cooling architectures.]

[Any new hardware must be recalibrated for thermal resistance, flow rate, response latency, and safety margin.]

The fifty-one point two percent could not be crammed into promotional pages as a universal energy-saving rate.

Change a batch of boards, change a type of cold plate, or even shorten the hoses by a few meters, and the original parameters could all become invalid.

After unlocking the hidden window, the twenty-seven 170HX cards possessed more available VRAM.

After connecting to the MPS Thermal Fabric, they formed for the first time a computing cluster that could continue working amidst cooling failures.

Jiang Lin created the Real World achievement package.

[MPS-ThermalFabric_Reality_v0.1]

The achievement package was split into six parts.

TF-Runtime is responsible for the joint scheduling of computing tasks and cooling branches.

TF-BranchLedger records continuous heat dissipation capacity, response window thermal budget, in-transit heat, and recovery time.

TF-CoolingReady_API feeds the cooling readiness status into the task migration chain.

TF-CabinetController manages temperature, pressure, flow rate, valve groups, and pump groups.

TF-FaultTest stores specifications for flow interruption, leakage, sensor drift, pump group degradation, and bypass takeover.

TF-CommissioningKit handles the thermal resistance, flow rate, and response latency calibration during new cabinet integration.

The material parameters, equipment numbers, maintenance port structure, and original service protocol of TM-7 were all left behind in the Wasteland sealed area.

[RING-3_LOCAL_THERMAL_SERVICE_PACKAGE]

[Permission hierarchy: Absolutely sealed]

Inside the Real World achievement package were only things achievable with standard copper cold plates, ordinary pump groups, solenoid valves, plate heat exchangers, dry coolers, and industrial sensors.

The future engineer had left behind a set of ideas.

From the summer of the twenty-fourth year to May of the twenty-sixth year, Jiang Lin spent nearly two years rewriting it into the pipes, valves, copper plates, and code of 2022.

Afterwards, he opened the Real World translation planning page.

[Low-Entropy Heterogeneous Computing Center Phase I]

[Construction location: Jiangcheng Main Machine Room]

[Beijing R&D Center: Retain verification node]

[Phase I target cabinet entry: 96 170HX-class boards that passed stability screening]

[Target schedulable VRAM: No less than 3TiB]

[Cooling architecture: MPS-ThermalFabric]

[Scheduling architecture: MPS-Scheduler]

[Fault ledger: MPS-FaultLedger]

[Evidence boundary: MPS-EvidenceGate]

The Real World procurement list unfolded downward accordingly.

Servers, cabinets, power distribution cabinets, pump groups, cold plates, sensors, plate heat exchangers, outdoor dry coolers, energy storage, and backup power supplies.

This already exceeded Zijing Apartment and the capacity of the test rooms at the Beijing R&D Center.

A large number of servers would bring noise, heat, power load, fire protection requirements, and 24-hour maintenance.

Beijing was suitable for retaining small-scale development, calibration, and fault reproduction nodes, while the real main machine room should be placed back in Jiangcheng, close to the manufacturing, procurement, and long-term O&M system of the Low Entropy Workshop.

There was already an answer to which route the computing equipment should take.

There was also a first version of the answer on how tasks would remain after a local cooling failure occurred.

On May 22nd of the twenty-sixth year, Jiang Lin completed the triple verification of the code, blueprints, test records, and failure ledgers.

The RING-3 original service package continued to be sealed in the Outpost storage array, without being written to any carrier prepared to return to the Real World.

Two hard drives brought from the Real World and specially reserved for the return only saved code, blueprints, test records, and failure ledgers re-verified through Real World materials, Real World components, and Real World interfaces.

Four servers exited the computing state one by one.

The circulation pump continued to turn until the residual heat in the cold plates, hoses, and buffer tanks completely receded.

Afterwards, the pump groups, valves, and servers were powered off, and the liquid cooling loop remained sealed.

The A100 remained at the baseline node.

Twenty-seven 170HX cards remained in the four groups of servers, while another five backup cards were re-packaged and put into the anti-static cabinet of Workshop No. 2.

Four servers, two sets of storage, and the entire liquid cooling system were also all left behind at the Outpost.

Twenty-six years of storage, startup and shutdown, disassembly, assembly, and maintenance had consumed most of the remaining lifespan of these devices.

Bringing them back to 2022 had little value; leaving them at the Outpost could at least continue to shoulder verification, computation, and spare parts.

The only thing Jiang Lin finally connected to the main connection band was these two return hard drives that had completed image verification.

The OR-MAINT-NE73 communication link dropped to its minimum maintenance power.

The TM-7 shell maintenance port was still buried underground in Northeast Seventy-Three.

That shell maintenance door had already been opened, and Jiang Lin did not submit permission requests any deeper.

The twelfth Wasteland task list unfolded again.

[Restore and expand Outpost power supply, ventilation, and storage environment: Completed]

[Establish A100 baseline node: Completed]

[Confirm the hierarchy where 170HX limitations reside: Completed]

[Acquire long-term operational computing nodes: Completed]

[Advance MPS-Agent α Real World adaptation: Completed]

[Enter TM-7 shell maintenance port: Completed]

Below the list were the four core achievements left behind over these twenty-six years.

[170HX hidden resource runtime: Completed]

[One-TiB-class heterogeneous near-storage computing cluster: Completed]

[MPS-EvidenceGate Real World adaptation: Completed]

[MPS-ThermalFabric Real World dimensional reduction: Completed]

The workshop gradually quieted down, leaving only the Wasteland storage array and low-power monitoring nodes running.

The fans were still rotating at low speed, and with every circle the blades turned, they sent a small segment of wind sound into the Outpost.

Twenty-six years ago, when the four servers had just landed here, there wasn't even a single liquid cooling pipe in Workshop No. 2.

Today, the four servers were still lined up along the wall. The chassis and interfaces bore numbers from repeated disassembly and assembly, and beside the wall were four extra branches, two welded buffer tanks, rows of measurement points, and tags filled with failure numbers.

These devices and traces would all remain in the Wasteland.

All Jiang Lin took away were two hard drives, along with the knowledge, judgments, and twenty-six years of failure records already left in his brain.

He grasped the main connection band.

The sound of the fans cut off.

The lights of the Beijing R&D Center fell into view, and the electronic clock on the wall had just ticked past one grid.

[September 30, 2022]

[06:01:00]

In the hallway outside, the cleaning staff pushed a cart past the corner, the sound of the rollers going from close to far.

At 8:30 this morning, the Low Entropy Workshop was going to hold its first batch delivery production preparation meeting.

At 2:30 in the afternoon, the artificial intelligence discussion in Nanjing was still waiting for him to connect online.

Jiang Lin sat in front of the computer and plugged the hard drive into the offline verification terminal.

The two images passed sequentially.

New document.

[Low-Entropy Heterogeneous Computing Center Phase I Procurement and Construction Task Statement]

First line: Board procurement.

Second line: Servers and cabinets.

Third line: Power supply and distribution.

Fourth line: Liquid cooling branches and end heat exchange.

Fifth line: Jiangcheng main machine room site selection conditions.

The cursor fell to the sixth line.

Jiang Lin wrote down the primary acceptance indicator.

[After any isolatable main branch suffers a flow interruption, already-started tasks can still safely remain.]

The twenty-six years in the Wasteland settled their accounts here.

One minute in the Real World had just ended.

The computing center of the Low Entropy Workshop started from the sixth line.

Prev Next

🔊 Text To Speech

Listen while reading

Ready