Saurabh
"As machines make discovery faster, people may come to see theoreticians as extraneous, superfluous, and hopelessly behind the times. Knowledge about a particular area will be less treasured than expertise in the creation of machine-learning models that produce answers on that subject."

In the summer of 2002, Dr. Devendra Jalihal of the Electrical Engineering Department at IIT Madras, gave me sixty diskless PCs. I set up about assembling them at the Integrated Electronics Lab in the basement of the department building. By day, these computers served as a computing facility for undergraduate students. By night, they turned into a compute cluster.

Compute clusters or now fairly common. In those days they were still fairly novel. They were called Beowulf clusters, and were used almost exclusively by the scientific community, to replace their supercomputers with something more flexible and affordable. I remember fixing Linux kernel drivers for the cheap NE2000 knock off cards supplied with the PCs, and burning EPROMs using an external programmer.

With the compute cluster built, I set about looking for problems to solve. I has been taking courses in operations research from Dr. Srinivasan G. Perhaps the most patient and loving teacher I have come across, I invariably slept through his classes as he laborisously walked our thick heads through OR algorithms. I asked him for a good problem to solve, and he connected me to Gowri Krishnamurthy.

As part of her undergraduate research, Gowri has been studying the manufacturing processes at MRF's factories. She wanted to schedule jobs for their various machines "just in time", so that elaborate planning was not required and the factory was more flexible to changes in demand.

I first rigged our cluster to run LAM MPI. I then wrote a C program that used genetic algorithms to solve Gowri's scheduling problem. The program would always converge within seconds at the optimal solution. We realised that a cluster was an overkill for a real world problem of this size.

In any case, results were found and documented. We eagerly went to meet Dr. Srinivasan. In his affable, genial manner, he carefully explained to us how our work wasn't a work of "genius" in any way. We had no idea why the results we had were optimal. We had, in fact, no proof that they were the optimal solutions to the problems at hand.

Running a genetic algorithm to find a minima or maxima was one thing for a cluster. Running through every possible solution to verify that the claimed minima or maxima was indeed to the actual minima or maxima, was entirely another. Our cluster would take days to verify each combination. Which meant that every morning I would need to save my work before the day's classes, and restart the job at the end of the work day.

What Dr. Srinivasan tried to tell us was that an algorithm that generates the ideal solution, and a mathematical proof of its optimality, was really what we should aim for. I am not sure if Gowri ever found such a proof. I had my own fears, which SriniG was only exacerbating. My research work was part of my Master's course in Communication Systems. Almost all of my peers were doing theoretical research, writing novel mathematical proofs (or so they claimed). I was one of the rare few doing engineering and actually building something. It took multiple discussions with Dr. Jalihal to assuage my fears that doing engineering and building something was a perfectly acceptable research work for an engineering student. Ah, youth :-)

There were other clusters that I helped build. And there were other problems that I solved on the cluster at IEL. One was a problem given by my control system's professor, to model the human visual cortex. Again, a purely machine learning algorithm, it was able to correctly identify lines and boxes quite easily. No one had proof for why the software worked, but were indulging in biomimicry. As long as the software did what the human brain does, we were fine.

And now we are in the age of AI and ML, where no one is entirely sure why stuff works the way it does. Sweatshops across the world run by Amazon, Google and Microsoft employe young people by the thousands, who listen into our conversations and flag them to be fed as good and bad samples into the behemoth. Since we don't quite understand the inner workings of these engines, we feed them with more and more data in the hope that they get smarter and more accurate. But there really never is enough data, in a world inhabited by 7 billion people, millions of languages and cultures, and a million billion stories.