Warning: session_start(): open(/opt/alt/php85/var/lib/php/session/sess_2aa511af9a28dad3b0e026b36647465f, O_RDWR) failed: Disk quota exceeded (122) in /home/u377687657/domains/novelfull.in/public_html/db.php on line 31

Warning: session_start(): Failed to read session data: files (path: /opt/alt/php85/var/lib/php/session) in /home/u377687657/domains/novelfull.in/public_html/db.php on line 31

Warning: session_start(): open(/opt/alt/php85/var/lib/php/session/sess_2aa511af9a28dad3b0e026b36647465f, O_RDWR) failed: Disk quota exceeded (122) in /home/u377687657/domains/novelfull.in/public_html/header.php on line 3

Warning: session_start(): Failed to read session data: files (path: /opt/alt/php85/var/lib/php/session) in /home/u377687657/domains/novelfull.in/public_html/header.php on line 3
This Top Student's Vast Amount of Knowledge Chapter 89 - 89: Chapter 89 Seventy Million | NovelFull
Reading Settings
Font Size
16px
Line Spacing
1.6
Reading Width
900px
Font Share
Theme
Text To Speech

89: Chapter 89 Seventy Million Little Stones

In front of the whiteboard, Yin Hang was still staring at the blind box disassembly problem.

His brain was running at high speed, trying to find even the slightest flaw in this intricate probability network.

Yao Siyu stood on the other side, holding a blackboard eraser in her hand, having already ruthlessly crossed out the old recurrence formula.

Jiang Lin stood aside, but did not continue explaining further.

These few problems were already enough at this point.

The true value of the preliminary questions for the Alibaba Global Mathematics Competition has never been about feeding the final bare answer to others, but about the process of seeing that line.

First, abstract.

Strip away all the fancy commercial attire from the extremely complex blind box probabilities and card pool pity mechanisms in the Real World, turning them into a pure Markov chain.

Then, compress.

Compress tens of thousands of node states based on symmetry into a minimal one-dimensional array related only to the remaining uncollected categories.

Finally, verify.

Use rigorous logic to prove that this dimensionality reduction compression did not lose necessary information at any tiny probability bifurcation point.

The marker in Yin Hang's hand spun half a circle on his fingertips, and the cap tapped against his palm with a snap.

He turned his head, looking at Jiang Lin with a gaze akin to looking at a monster, and suddenly asked, "Don't tell me that all eleven of these problems were finished just like moments ago, seeing through the underlying structure at a glance, and then writing them down without any stuttering?"

Jiang Lin thought for a moment, and shook his head earnestly: "Not entirely. One algebra transformation problem leaned more towards pure calculation and typesetting expression, lacking such a strong sense of structure, and was purely piled up by computing power. Moreover, I also tried a few wrong paths."

Hearing the words "wrong paths", Yin Hang's shoulders slumped abruptly, and he seemed to let out a long sigh of relief: "Damn, so you also take wrong paths. I thought you had a built-in quantum computer in your brain."

"The wrong paths I've walked might be more than any of you."

As Jiang Lin spoke, he even gave an example, recounting how he was once confused by the complex appearance of a number theory problem and walked into a dead end.

"So, finding the correct route and eliminating the wrong directions took you about forty minutes. Afterwards, organizing the submitted version with LaTeX took over an hour."

Yin Hang was amused and exasperated by Jiang Lin's solemn complaints.

"Thank you, you finally admit that you spent time and finally act like a human. But just finding the route for this blind box problem took me four times as much time as yours, and I still haven't found all of it."

The atmosphere in the room finally relaxed completely.

The air, tense due to high-intensity intellectual confrontation, vanished like smoke with Yin Hang's self-deprecating remark.

Yao Siyu couldn't help laughing either: "So Director Lu telling you to treat B304 like your own home isn't without reason. With you here, at least you can let us know to what extent a problem can be torn down, so we won't think we've seen the ceiling just as we touch the threshold."

Seeing Jiang Lin starting to pack his things and knowing he was about to go back, Meng Che on the side handed over a few downloaded data descriptions and printed paper drafts, and asked casually: "What else have you been busy with besides the Alibaba preliminary rounds? It feels like you've been elusive all month."

"Writing a data checking tool."

"What is it for?"

"Quant-related, I suppose."

"You even do quant?" Meng Che said in surprise.

He was researching machine learning and knew very well how deep the waters in the quantitative finance field were.

"How old are you, and you're already doing quant?" Yin Hang also had an incredulous look, "Don't tell me you've already started trading stocks."

"Just taking a casual look, helping someone do some basic data sorting."

But the B304 trio obviously didn't believe it.

A person who could finish the eleven problems of the Alibaba preliminary rounds like chopping vegetables—when they said "just taking a casual look", it was definitely not as simple as drawing a few K-line charts in Excel.

However, seeing that Jiang Lin had no intention of explaining further, the three of them sensibly didn't ask more.

Geniuses always had their own secret domains; they still had this bit of tacit understanding.

In the evening, Jiang Lin returned home accompanied by the sunset glow of Jiangcheng.

After dinner, he chatted idly with his parents for a while before returning to his bedroom to attend to his own matters.

The workstation computer was turned on, and there was a new email.

Sender: Shen Chengye.

Title: QF-OLDLIB-001 First batch of desensitized materials uploaded.

The body text had only a brief sentence: Jiang Lin, the research process audit of the small private equity old factor library officially begins, resources have been opened.

Jiang Lin clicked the link and began downloading the materials.

Compared with the tick-level high-frequency market data of a dozen GB or more that he had processed before, the files were not too large, and the compressed package was less than 3 GB.

But in the context of quantitative research, the information density of this volume was terrifying.

After decompression, the directory tree expanded.

Several massive metadata tables.

Several backtest configuration summaries.

A desensitized factor output matrix.

A very crude historical notes record.

And an extremely crucial status table.

Jiang Lin clicked open the status table. The table densely listed 437 factor IDs.

In the status column:

Some were marked with: active (factors currently still providing signals in live trading).

Some were marked with: deprecated (factors eliminated due to performance decay).

Some were marked with: merged (factors merged due to excessively high correlation with other factors).

Some were marked with: deleted (factors that were deleted).

There was also a batch shockingly marked with: unknown (unknown status, even the researchers themselves didn't know whether this thing was still running).

Jiang Lin's gaze lingered on the two words deleted and unknown for a long time, and his thoughts suddenly drifted away.

In any quantitative institution, or even any scientific research system, successful things would always be remembered by someone.

They would be written into PPTs, hung in annual reports, and used to boast to investors.

However, failed things were often the first to be deleted.

To cover up the embarrassment of having no achievements for several months, or to make the codebase look clean, researchers would press the Delete key without hesitation.

And a research system without failure records was most likely to mistake survivors for truth.

In the financial market, this was called survivorship bias.

When you only saw those successful strategies, you would feel that the market was full of rules.

But if you couldn't see the 99% of discarded factors that suffered heavy losses in live trading due to overfitting, you would repeat the same mistakes in your next research.

Jiang Lin took a deep breath, skillfully opened the terminal, and created a new project folder.

Then he opened audit_log.md with Vim and typed out the supreme guiding principles for this audit project.

First line: The old factor library is not a treasure trove in the first place, but a cemetery of failure records.

Second line: This project does not look for god factors, does not predict the future, but only looks for "how the research process deceived itself".

Third line: All conclusions must be strictly bound to four dimensions: data version, sample pool version, factor version, and backtest configuration version. Discussing performance detached from versions shall be treated as academic fraud.

After writing these three lines, he officially began the first round of code-level scanning.

The logic of the initial audit script was not complicated.

Using Python with Pandas and Dask, Jiang Lin wrote several guard components.

Uniqueness check: Check whether the factor IDs are unique, whether version numbers are continuous, and whether each factor is strictly bound to the cleaned data version used at that time.

Sample pool drift: Is each backtest result bound to a definite sample pool definition?

Failure record completeness: Do the failed factors retain environment slices of detailed deletion reasons and failure dates?

Homology check: Do identically named or similarly named factors repeatedly appear?

Friction cost check: Have transaction cost assumptions artificially drifted across different versions?

Enter, run.

Fifteen minutes later, the first round of scanning results came out. They were ugly, shockingly ugly.

Among the 437 factors, more than 80 lacked complete data version number constraints. This meant that if the backtest were re-run now, it would be entirely unknown what data was used to run it back then.

27 factor versions were fractured, jumping directly from v1.2 to v3.0; what happened in between was known to no one.

13 factors had completely different names, but the mathematical logic in the notes was highly similar.

There were also a few factors whose names looked like three completely different directions.

But after Jiang Lin's script calculated the cross-sectional correlation for their desensitized output matrices, it found that their Pearson correlation coefficient was unexpectedly as high as 0.98.

This showed that this was simply the same idea, repeatedly copied, with parameters slightly tuned, renamed, and re-run amid performance pressure and process out-of-control, trying to bump out a better-looking Sharpe ratio, ultimately leaving three extremely similar shadows in history.

This was a pure data mining disaster.

Jiang Lin expressionlessly continued to run the second round: Factor output version sensitivity test.

Round three: Sample pool time series drift test.

Round four: Failure factor record completeness penetration.

As the tasks grew heavier, the small chassis fan under the workstation began to hum continuously, emitting a dull buzzing sound.

Jiang Lin glanced at the system monitoring on the secondary screen.

Temperature: 68 °C, normal.

Hard drive read/write: normal.

Memory usage: normal.

The CPU usage rate, however, repeatedly shot up to 100% during the execution of a few specific scripts, blindingly red.

The program didn't freeze, and the progress bar in the terminal was still moving.

But it was slower than Jiang Lin expected, unreasonably slow.

For most data scientists or quantitative researchers in the Real World, encountering this situation, the first reaction would definitely be that the machine wasn't powerful enough.

"Boss, we need to buy a better CPU and switch to AMD Threadripper."

"We need to add more memory to cram all the data into memory."

"Go to AWS cloud server and spin up a 124-core instance to run in parallel."

Using hardware brutality to cover up software inefficiency was a common ailment in peacetime eras and resource-abundant environments.

But this wasn't Jiang Lin's first reaction.

Deep in his mind, the countdown belonging to the Wasteland World always existed.

In the resource-depleted Wasteland, there was never an option to buy another one.

There, when the machine was slow, you couldn't ask for resources first; you had to ask first: Why was it slow? What was eating up the computing power?

Jiang Lin decisively pressed Ctrl + C, stopping the subsequent scanning tasks.

Open the Python performance analysis tool, re-wrap the extremely slow-running code from just now with cProfile, and then connect it to SnakeViz for visual analysis.

Twenty minutes later, a detailed call stack flame graph appeared on the screen.

The true bottleneck surfaced.

The monster that swallowed the most wall-clock time in the call stack was not those large matrix multiplications of millions of rows.

It was not the complex HDF5 file reading, nor the front-end rendered chart generation, nor some metaphysical machine learning complex model.

Instead, it was a few actions so small they were almost inconspicuous: sorting, ranking, and bucketing.

In quantitative backtesting, it is often necessary to neutralize and sort the factor exposure values of stocks within specific industries and specific market capitalization ranges.

The number of stocks participating in sorting each time may not be many, only five, eight, or sixteen.

Viewed individually, sorting five numbers, regardless of the algorithm used, was every time as fast as costing nothing, not even needing a millisecond.

But the problem lay in the multiplier effect.

There were more than four hundred factors in QF-OLDLIB-001.

Three years of historical versions.

Multiple dynamically adjusted sample pools.

Various backtesting configurations.

Every day, every industry, every state tag, and every data version required slicing out layers of cross-sections to perform grouped ranking.

Consequently, these tiny little actions were nested within a massive loop network and called repeatedly.

Jiang Lin printed out the call count of the function that consumed the most time in the call stack.

Seventy-eight million four hundred twenty-one thousand nine hundred and six times.

Looking at this astronomical figure, he was silent for over ten seconds, and then in the project's audit_log.md, he typed out the following passage:

"What truly slows down the entire complex system is never the occasionally appearing great mountains, but the small stones that must be moved seventy million times every day."

After writing this sentence, he paused for a moment.

To make it understandable even for mediocre engineers who might take over this audit report in the future, he added another more colloquial explanation.

Sorting five test papers from high to low scores at once is not difficult for anyone.

What is difficult is that the system requires you to repeat the action of sorting five test papers seventy million times within a single day.

The sorting algorithms in the standard library have been proven in computer science to be extremely excellent.

But the premise of their excellence is being general-purpose.

The standard library is like a large automated logistics sorting center covering tens of thousands of square meters.

It can handle ten thousand test papers.

It can handle one million packages.

It can handle complex data structures with all sorts of strange objects.

It is powerful, general-purpose, and absolutely reliable.

But if on your assembly line, what is delivered each time is forever only five small packages, and they must be delivered seventy million times a day,

then every single time you start up that hugely power-consuming large logistics center, letting the conveyor belt idle, letting the robotic arm seek addresses, and executing the massive sorting logic,

that is an unforgivable waste.

From the perspective of low-level code, this waste manifests as the complex function call overhead retained for generality.

Dynamic type checks performed to handle polymorphism.

To accommodate different array lengths, the standard process retains a large number of conditional branches.

And once these branches are repeatedly triggered within hot loops, they slow down the instruction pipeline that modern CPUs rely on most.

In fact, it is not that the system cannot sort.

Rather, the process is too heavy.

So heavy that every clock cycle of the CPU is being consumed by meaningless management logic.

What Jiang Lin needed to do now was not to overturn the classic sorting theory written by Donald Knuth in The Art of Computer Programming, nor was it to invent some world-shaking new algorithm.

He just needed a set of fixed gestures.

Five test papers.

Look at the first and second ones.

Whichever is larger goes in front; swap them if necessary.

Look at the third and fourth ones again.

Swap if necessary.

After a few extremely fixed comparisons, the order naturally comes out.

Do not ask redundant type questions.

Do not open redundant memory allocation processes.

Do not prepare any redundant boundary checking tools for that one million test papers that do not exist at all.

There are no data-dependent loops, no runtime temporary path selection.

The comparison order is nailed down before compilation, and what remains is merely comparisons and swaps between fixed positions.

Only process the numbers at these five positions.

This is the core secret of optimization for small-scale data with extremely explicit boundaries in the field of high-performance computing.

At 1:20 AM, with absolute silence all around, Jiang Lin wrote the signature of the first version of the function in the newly created C-language extension file.

The function name was very ugly, not even looking like something in an elegant algorithm library.

rank5_fixed_v0.

It did not attempt to sort all arrays in the world, only processing five float64 factor exposure values, five validity tags, and five original position numbers.

What it output was also not a pretty new array, but a set of business rankings and a set of masks.

It was like a weird wrench in a Wasteland workshop, whose handle was forcibly welded bent in order to tighten a specific screw on the chassis of a certain model of engine.

But what Jiang Lin needed right now was precisely this single-minded and brutal wrench.

After finishing the first version, Jiang Lin did not rush to replace it directly into the Python audit main process.

As a survivor in the Wasteland who had witnessed a single decimal point error cause an entire proof to fall short at the last hurdle, he had a morbid rigor regarding replacing underlying logic.

He first made a baseline.

Whatever the original process outputted, the new function had to output the exact same thing.

In quantitative finance data, reality is always dirtier than theory.

No duplicate values (ideal state).

Duplicate values (two stocks have completely identical factor scores).

Missing values (NaN, a certain stock was suspended from trading today with no data).

Extreme values (Infinity).

Negative values.

Equal values while needing to maintain the original relative order (stable sorting requirement).

Every single case had to be aligned.

When there were duplicate values, whoever was in front in the original array must still be in front after sorting.

When encountering NaN missing values, the quantitative rule was not mathematically treating it as the maximum or minimum, but rather having to filter it out individually according to the project's preset rules and tag it with MASK_NAN, letting the remaining numbers continue to sort.

This was not sorting in the pure mathematical definition, but ranking rules within a financial audit project carrying strong business attributes.

The two must never be confused.

At 2:30 AM, Jiang Lin rubbed his sore eyes.

The first version v0 passed unit tests containing twenty thousand edge test cases.

The results were entirely correct.

The speed had improved, but not much, only about 15% faster.

The first version was only to verify that this strongly coupled direction was viable.

Jiang Lin was not surprised.

He reopened the C code.

And began to truly squeeze out the performance.

Delete all unnecessary state judgments.

Completely unroll the only bit of loops into straight-line code.

Completely weld shut the possibilities of various data types that might appear, limiting them to the memory layout that would actually be input into this project.

At 3:18 AM, the second version came out.

He did not expose this function as a small toy called iteratively by Python loops.

That seventy million cross-language calls would in itself become a new disaster.

What he truly wrote was a batch processing entry point.

Receiving millions of 5-tuples in contiguous memory all at once, running the entire fixed sorting network internally at the C layer, and then spitting the ranking matrix back to Python.

Running the tests again.

The results were consistent.

The speed increased to 30%.

But this was still not enough.

Jiang Lin's brows furrowed slightly; he felt there was still excess fat in the code.

He took out a blank sheet of paper and a pen from the drawer.

On the paper, he drew five circles, labeling them with serial numbers: 0, 1, 2, 3, 4.

Then he began drawing lines between the circles.

What he was writing now was not code, but actions.

In low-level assembly instructions, compare-and-swap is an extremely cheap action.

As long as there are no prediction failures caused by if branches, instructions can flow smoothly through the CPU pipeline like water.

Compare 0 and 1 (the larger goes to the right).

Compare 3 and 4.

Compare 2 and 4.

Compare 2 and 3.

Compare 1 and 4.

Compare 0 and 3.

...

Every step was like performing a precise mechanical manual adjustment.

For five numbers to be sorted, there is simply no need for the program to think about what to do next every time it runs.

The route can be nailed down in advance.

Just like water flow passing through pre-dug maze ditches, regardless of the water volume, it will eventually flow out from the designated exit in order of size.

As long as on this grid route, all possible $5! = 120$ types of initial arrangements can eventually be correctly channeled into an ordered state, that is enough.

This is an extremely niche yet extremely hardcore concept in computer science.

Sorting network.

It is not clever at all.

It is completely helpless facing 1,000 numbers.

But it is very stable.

It is not general-purpose, but it is the ultimate killer weapon tailor-made for high-frequency, small-scale tasks.

Writing up to here, Jiang Lin stopped his pen.

Suddenly he remembered the magnetic geometric Rubik's cube placed at the corner of the desk in B304 this morning, and he remembered that blind box question from the Alibaba Mathematics Competition.

He remembered what he had said to Yin Hang: Do not let the patterns lead you by the nose first; find the upper bound first, then find the equality-achieving construction.

Optimizing the underlying sorting logic is the same.

Do not be scared by the big word "sorting" that has been talked to death in textbooks first.

What he needed to deal with now was not sorting science at all, but simply the minimum sufficient set of a limited number of comparison, swap, and combination operations among five memory locations.

This was a very small world.

So small that all its states could be mathematically exhausted.

So small that it could be rigorously proven.

But it was important to the extent that as long as this action was repeated seventy million times, it would become a structural bottleneck dragging down the massive financial system.

At 4:10 AM, the skyline outside the window already faintly had a hint of cyan-gray.

Jiang Lin typed out the final set of instructions; the third version was complete.

Prev Next

🔊 Text To Speech

Listen while reading

Ready

Warning: session_start(): Session cannot be started after headers have already been sent (sent from /home/u377687657/domains/novelfull.in/public_html/header.php on line 80) in /home/u377687657/domains/novelfull.in/public_html/footer.php on line 3