giovedì, giugno 21, 2018

White King, Red Queen: a data based sidetrack. Part V.


Our journey has brought us from an original quote from Hofstadter, through Faust and trees until the final victory of machines against humans.
We have repeatedly seen how evolutionary pressure for economic survival, academic prestige and generic coolness pushed further and further on the path of chess automata. We started in a world where it was unthinkable that chess could be played in some acceptable manner by machines, and only 30 years later, we got to know a world in which the world champions would be dominated by smartphones. This took 30 years, some great minds and the direct intervention of the Red Queen, making sure that everyone would be forced to run and run and run, without a single actor being able to hold the crown for more than some years.
Our journey would be incomplete, though, if we would not mention the living fossil, the indomitable crocodile of the digital chess arena: the ecosystem created by ChessBase. So we will shortly sidestep the world of chess engines and chess playing and give a short look to the world of chess databases.

Enter Chessbase

What is ChessBase? It is a German GmbH, based in Hamburg, which sells world-wide mostly one, single digital product and related ecosystem material. This product, a database software and related game collections and training material, can be downloaded or ordered per postal service. In spite of the apparent obsolescence of this business model, it exists since 1985 and features a rock solid € 100 t net income and € 2 m equity1
What they sell exactly, and how did they manage to survive for so long?

Garry, again

If you can read German or possess a copy of Kasparov's Child of Change: An Autobiography (1987, Hutchinson), you will be able to revive the exact moment in which Garry was confronted with the first prototype of the ChessBase software. He describe the festive Christmas atmosphere in a home in Switzerland in which he expressed his wish for a database of games, its later experience with the first version of ChessBase, and how he did use this tool to its advantage in simultaneous matches. This was exactly the period in which Garry was fighting for the world championship against another Russian titan, Anatoly Karpov. Understandably, this boosted the acceptance in the chess community and made possible to expand rapidly and reach widespread diffusion. So they became the leading vendor of chess databases.

Inertia can kill, but must not

This was 1985. If you remember, this was more or less the time in which chess engines were spreading and becoming dominant: H + G, this other German company we met here was making money, too, with a computer based chess product. They were selling hardware-based engines, true, but they had the know-how, they could have moved to PC-based engines and survived. Why did they miss it? We will never get a definite answer. Fact is, ChessBase recognized that they have to diversify their product. They started offering only generic game collections. The more games, the better. They then started to offer specifically tailored and reviewed game collections, typically focused on a single opening. They somehow recognized that the client used to browse and search the game collections could be improved on and used as a training tool. They started selling training material. With time, they became more a publisher of chess material than a seller of a digital product. Chess players are a conservative folk, and often recruit from persons with higher economic status than the average: if they have a product which is good and reasonably priced and serve the goal of comfortably analyzing games and preparing for tournaments, they will not switch to a random competitor. In this way, they could bloom and flourish for more of 30 years.
The Red Queen, however, may tolerate some hesitation, but clouds are already on the horizon, and who does not run will get beheaded faster than he can say "Check!".

The next menace

Using concepts from mapping and co-evolution, it is not difficult to understand what the next threat for ChessBase will be (Mr Wüllenweber, free advice for you here!): the commoditization of their business model2. Commoditization in digital landscapes nowadays almost always happens in form of cloud services or open source, or publishing platforms. What could happen here?
  • Commoditization of analysis: game collection platforms expanding in the analysis region. ChessBase is inherently a home PC product. You have the client, you use your local chess engine to analyze a game, give hints and similar. Given the extremely low cost of cloud-based computing resources, web-based competitors can offer analysis solutions. Indeed, Chessgames, a historical web page offering game collections (more of less a web-based version of ChessBase), offers online analysis on a Stockfish server, too.
  • Commoditization of training: Online game servers offering publisher platforms for chess blogger and freelancers. This is more or less the business model of chess.com. You register for playing, you get some articles and training for free. But if you want the full training experience, you have to pay a subscription.
  • The doomsday device: open source. Stockfish, for instance started some weeks ago publishing high quality training games on google drive. Given that these games exceed grandmaster level by a huge amount and the collection is growing at alarming pace, it will put a big pressure on pure game collection vendors. What will come next is in the stars.
Mr Wüllenweber, be warned!

1you can get all basics financial information about GmbH and AG German companies from the Bundesanzeiger
2 I plan to eventually publish a map about this in a reprint of this article, so stay tuned! 

 

domenica, giugno 17, 2018

Football learns from chess and finally uses Elo!



Finally, we are there after 50 years: the FIFA introduces the Elo rating as a measure of national teams relative strength!

For the chess fans among us, a well-deserved recognition for a classical tool from the chess world. For for the serious football fans a huge opportunity to improve understanding of playing strength and to become more informed about the real performance of the beloved team.

Let us start our journey with the hero of this story Arpad Elo.

Arpad Elo

Let us clear this in the very beginning: Elo is a name, not an acronym. Specifically, Elo is the family name of an US American theoretical physicist and chess player of Hungarian origins, who, in the 1950es, invented a statistics-based system for rating the relative strength of chess players.

As a theoretical physicists is difficult to find traces of him. He seems to have done some research in analytical chemistry, but his activity was clearly focused in chess. Being a strong chess player on his own, he was involved in the United States Chess Federation (USCF) from the very beginning and convinced both USCF and then the international chess federation (FIDE) to adopt his system.
He died in 1992, aged 89.
He has a good wikipedia page, which is as usual a very interesting reading.
By the way, from this short introduction, it should be clear that Elo should be written like Elo, and not like ELO.

How does it work?

The idea behind the Elo-rating is to have for every chess player a number, the rating, with the following property: if we have to players A and B, with ratings r(A) and r(B), the average results of their games should be a function of the rating difference alone. This is very different from other types of ratings. Imagine, for instance, to use the season rankings from a typical football league as your rating. These are calculated by giving each team 3 points on a win, 1 point on a draw, 0 points on a loss. Since these numbers increase with the advancing season, so do their differences, although the relative strengths of the teams do not change (in first approximation).
In practice, the are several different implementations of Elo systems. They usually use some variation of following procedure: when two players have a match, one computes their rating difference, say D, and divide this by a scaling coefficient S. This is the expected result of the game. This is done for each of the two players, so each player will have an expected result. For chess the expected results will sum up to 1. This is because you get 1 point for a win, half for a draw and none for a loss. So always 1 point is awarded. After the game, one compares the expected results with the actual result for each player. This difference is multiplied with another number K and then added to the rating of the player.
There are many variations on this implementation. As usual, every rule calls for problems and exploits.

History of Elo 

Until the 90es the history of Elo rating has been quite straightforward. More countries would adopt it, more decisions in chess would be based on the Elo rating of the players. For instance, chess has titles, both on the national and international level. If you are a good player, you could achieve the title of National Master, awarded by your federation according to own rules: usually they include to have at least a certain Elo and to achieve a given Elo performance on 1 or more tournaments. What is Elo performance? Imagine you play a tournament: since Elo rating is about averages, you could average the Elo rating of the players you play against in the tournament, and consider your average performance in the tournament. If you played 6 games for 4.5 points, this would be 0.75. Comparing this results with the average rating (as we did before) yields the Elo performance for the tournament. Accordingly, FIDE awards the title of FIDE master, International master and International  Grandmaster based on similar criteria. World championship is, in contrast, awarded based on a lineal system.
Until the advent of internet, there was a problem: Elo would get updated rarely. In Italy, the country where I came from, every 6 months. Now, imagine that, for some circumstances, you would lose more Elo than usual in one term. Say, you got sick on the only two tournaments you played, performing way under your standard. You would then start the next term with a very low rating, making very easy for you to make gains by playing weaker players with the same rating as you. It is as if you would punch in a weight category lower than your natural one. If you now happen to play many tournaments in that term, you would be able to boost your rating way above your real strength. Interestingly, this has happened in 1997 in Italy, where Roberto Ricca could climb to the top of the Italian rating list in spite of being an average national master.
Modern online chess platforms, like chess.com updates your rating after each and every game, so this exploit is not possible anymore (just in case you would like to try).
Much research has been done in this topic: what if you want to estimate Elo from a static database of results? What if you want to estimate also the variance of the result and not only the average?

Elo in other disciplines

Football is not the first non-chess discipline to use Elo. Online based game-like platforms are a natural fit for this very easy system. Halo and Tinder seem to use it, too.
In fact, classical Elo implementation in chess is just a very simple data stream compatible machine learning algorithm for estimating the expected result of matches. Also shows that simple algorithms can go a very long way.
Interestingly, non official Elo-like ratings in football have been made available from freelancers since ages (like this), I remember to find one back in 2008, as I moved to computational neuroscience and tried to get an overview about machine learning and similar topics.
I hope you enjoyed this overview over the Elo system, if you would like to know more about Bayes Elo or have me dive into the difficulties of Elo for non-zero sum games like Football, let me know with a comment!

giovedì, giugno 14, 2018

Riemann vs Lebesgue (thank you Kia)

RandLintegralsFor decades now, I have wondered about the true difference between the Riemann and the Lebesgue integrals in analysis. The classical textbook explanation goes more or less like that: "Look, Riemann integrals are bad for commuting with taking the limit, let us try something else: if we subdivide the value space of the function instead of the interval of definition, everything turns well. So let us do it". I accepted this until some times ago. But then something happened.

Some insights from the mind of a know-it-all

I recently changed job. I like to have constructive conversations with my colleagues and I have worked hard at this in the last years. Unfortunately, I live with a heavy burden. There is a constant battle between the "know-it-all" (Kia) section of my brain and the Dr. Henry Jekyll part. She is continuously proposing to show how cool she happens to be and how much better would she run things and how it is possible that you don't see this and so on and so forth. In the course of time, Henry did indeed develop a strategy to cope with this: delayed gratification. He would not let Kia to start its rant immediately, but he generously would allow her a imaginary conversation with the culprit, impersonated by himself. There, Kia is allowed to use all the destructive rhetorics, cynicism, sarcasm and intellectual superiority she can muster. By the way, what does it matter that I changed job? It matters, since Kia get very very excited every time she does meet new persons, in particular in work. So I went through some hard times in the last months. Sometimes, however, Kia gives me insight.

If you are that good, explain me the difference between Riemann and Lebesgue integral

Impersonating a particularly nice and competent new colleague, short of nasty questions to throw at Kia, Henry just decided to resort to a random math question. By the way, note that my current job has nothing to do with academic maths, just to let you gauge the folly of my situation. Here he goes:

H: Well, if you are that good, give me the true difference between Riemann and Lebesgue integral.
K: How can you not know about this? I though you did know about math! Dear me, I will go through it, if this is really necessary.
H: Let me see...
K: Every small child knows that in Riemann integral you partition the definition interval and in Lebesgue integral you partition the range of the values. Already in his doctoral thesis, which I happen to have read, my dear friend, in contrast to you ...
H: Hang on. Why now it should be better to partition the range of values?
K: Ah! You see? You did not learn anything in high school! You don't have any meaningful convergence theorems for Riemann integrals, so you have...
H: Why?
K: Why what?
H: Why you don't have any meaningful convergence theorems for Rimann integral and why it gets better with Lebesgue?

Yes, why?

Functions with bounded variation

There they stopped arguing: it became clear to them at once that they did not really knew why the world is that way. So they started a more intimate conversation in candlelight. The first thing what occurred to them is that by subdividing the interval of definition in equal intervals, I could not really control how much the the true integral of the function on a subinterval will differ from the area of the local approximation of the integral. This is bad. For doing this, in fact, they should make the intervals small when the function changes much and large when the function changes little.

And then I just had an image in my mind of the cover of Riesz functional analysis book, which I happened to be reading in the library of Tübingen university in 2004. They dug deeper, and brought up the fact that there I learned there about functions of bounded variation. Ah! Bounded variation! After 14 years, this small piece of information started to make sense. After 14 years, I understood why functions of bounded variations are so important. After 14 year, I finally discovered that they help controlling how a function changes in a interval and this is of crucial importance when defining the Riemann integral. As this would not be implicit in their very definition, to hell with them!

In fact, functions with bounded variations can be integrated very well in the sense of Riemann, I discovered. Incredibly also this obscure Riemann-Stieltjes integral makes sense suddenly!

But we want the convergence theorems. Go down some lines in Wikipedia: the space of functions with bounded variations form a Banach space in its own right, if endowed with an appropriate norm! Oh my God, why nobody told me that 14 years ago! We actually can have convergence theorems for Riemann integrals. So, why Lebesgue Kia, why Lebesgue?

The problem is, to make the space of functions with bounded variation to a Banach space, the required norm is given by the total variation of the function plus its Lebesgue norm!

So, in order to have converge theorems in the sense of Riemann, you must already know of Lesbegue integrals! Ah! I will not sleep this night because of this I know it...

Thank you Kia!

domenica, giugno 10, 2018

White King, Red Queen: Machine over Human, part IV

They wanted to write chess programs (I); they did it and brought a chess program in every household (II); we also learned how to write our own chess program (III).

The next step would be to look into what has been repeatedly called one the turning points in the history of artificial intelligence: the two Garry Kasparov vs Deep Blue chess matches.

Indeed, the matches were extremely exciting, back then in 1996-1997. Kasparov was a true legend in the chess community. World champion at the age of 22 in 1985, he held the title until 2000. Already then he was widely recognized as a one of the greatest chess players of all time, featuring a aggressive, appealing playing style.

The challenger, Deep Blue, was a chess-specific supercomputer built by IBM both as research project and a publicity stunt.

The history can be summarized quickly: Deep Blue lost the first match in 1996 and won the revanche in 1997. Last year, Kasparov gave a great interview about that time.

In hindsight, it is clear that the defeat of the world-champion by a chess computer was a well predictable event: the question  should have been not whether but when in the decade between 1995 and 2005.

Strangely, it did not feel like that back in 1996. There were strong feelings about whether it would be really possible that a computer could possibly beat Kasparov. There were strong feelings about what would be the consequences of the defeat of humans by machines. Chess would be dead. Young people would'nt play anymore. And so forth and so on. The old "the world will go the dogs" kind of story.

Deep Blue won and nothing happened, apart of people talking about it for some weeks and then forget about it. Obviously, this would not stop computer scientists from working on the topic. Chess engines still were developed. In 2009, a mobile phone with Hiarcs 13 would win a super chess tournament and reach unprecented Elo ratings. Nowaday, grandmasters play chess engines only with odds.

And still, people would still play chess, new young talents will come to chess clubs, chess instructors will not go out of business. From my very individual perspective: the situation for chess players has improved over time since 1997.

One things changed, though: chess computers would be taken seriously from now on. The field of computer chess would become a competitive field in its own, played by its own rules and igniting a new set of innovations.

 

martedì, maggio 29, 2018

Never lose money (or anything else)

As I mentioned some days ago, I have just gone through a biography of Warren Buffet. One of his principles, it seems, has been to force company managers to always return sizable value to the shareholders, even when this means to reduce investments in growth. 

In a way, this reminds me of the agile obsession in software development to continuously return value to the stakeholder: small, stable, robust increments, delivering additional value to the company and the stakeholder. Abhor big-bang releases, make sure to get from a stable situation in another stable one and always make sure to give your shareholders some value, if only in the form of added knowledge.

I even suspect that the terminology used in the agile community is, at the very least, inspired from value-based investment theory. You often read about stakeholders, return on invest and similar vocabulary.

I have to admit that this point of view has a great appeal to me. I personally see applications of this principle in other situations. Some years ago, I was reading a book about relationships and communication in married couples (no, I cannot remind the title this time) and one of the behavioral psychologists authoring the book observed that much suffering in modern world arise from the underestimation of the effect of the pain arising from a divorce or otherwise separation and from the failure to objectively manage the unpleasant situation of a difficult relationship. The unexciting course of action would be to give to (and request from) the partner a small, but constant return on emotional invest. In the book, they even used the word "relationship account" and devised a structured approach to make (and request) small, continuous payments to this account. I have to admit that this structured approach has navigated my wife and myself out of dangerous waters in bad periods.

In contrast, a large investment with an uncertain return (separation and hope in a better future relationship, to stay in the current example) is often preferred to the small return of a current difficult relationship, and one fails to focus on effective measures to increase this small, but steady advantage.

As I was discussing exactly this topic with my wife some days ago (she always has to endure my rants about whatever book I read last), I just realized that this effect is wonderfully explained in in the masterpiece Thinking, fast and slow in terms of cognitive illusions. In fact, what happens here is the overconfidence in the own imagination of future situations due to the cognitive ease with which we ignore the difficulties which will arise in the future, just because we don't see them now. And, of course, we overestimate the effect of our actions since we can easily imagine them. 

A future great reward with little chance of realizing is preferred to less exciting options, in particular when the current situation is slightly adverse. So, never lose money or anything else!

venerdì, maggio 25, 2018

Reading about debt

I often talk to friends and family about books, often in a overenthusiastic manner. I like books and I like, more than everything else, to find useful stuff for real life in books.

If you're now expecting an apology of "Mindfulness for dummies" and the like, I have to upset you; I found out that the treasures are to be found in unexpected places.

As an example, take my last two reads: Debt, the first 5000 years and Buffett, the making of an American capitalist (not finished yet, but I also learned something very useful there, which does not have anything to do with investment, but let me talk about it some other time).

Does it sounds weird to have a leftist anti-capitalist book about debt in the direct neighborhood of the very generous biography of man with US$ 87 billions net worth? Don't worry about my mental stability, I do it by purpose. 

(Which of course is not a sign of mental stability, now that I think about it)

The first book investigates the origin of money in the credit instruments of ancient civilizations. Yes, it looks like it is an established fact that the Babylonian had financial instruments recorded on clay tables in a fictive currency (the bushel) which was only used for recording these financial transactions. And yes, they knew about compound interest. And no, they did not have any money you could exchange on the street. Crazy.

Anyway, there are also pages dedicated to the fact that the precise quantification of debts between neighbors and acquaintances tends automatically to depersonalize these relationships. This being the reason, for instance, for which we do not like to precisely keep track of who owes whom how much money if we go out with our best friend for a beer.

My wife is thinking since some time about turning her year long practice in mindfulness mediation and her teacher experience with it in a full-time activity mixing self-employment and community service. We talked a lot about financial independence and the like in the last weeks. And then BANG! I read this book and I understand why the financial dependence feels so threatening in a relationship. Suddenly, I have a framework on how to talk and think about it.

I think this is the sign of great books: you find insights and inspirations there which do not have anything to do with the apparent topic in the book.

Good reading!

lunedì, maggio 14, 2018

White King, Red Queen: a technical primer about trees, part III.

In the previous posts (I, II), we have seen the arms race around computer chess started. We have seen the first ideas about the possibility of building a true chess computer. We witnessed a German company doing it and being displaced by the newborn 486 and we left the an army of chess programmers getting ready to assault Gary Kasparov dominance over the world of chess.

At this time, we are talking about 1994, many people not directly involved in technology still wouldn't think that the end of human chess dominance was only 3 years away. I witnessed exactly this period as a young club player and I have vivid memories from that time. Indeed, the explanation I give you later is more or less what I got back then in 1996 by a chemistry professor interested in game theory and it still holds true.

The main reason for the supposed inability of computers, was that they were thought to purely use "brute force" to find the best moves, whereas master were known to be able to have long-term understanding of the game. The true skill of the masters, so the mainstream, was based on being able to recognize positional advantages which would bear fruit 20 moves or more later, and such a far horizon was thought to be unreachable by standard alpha-beta engines. Ooops, I did it. A technical word. I think the moment has come for you, my faithful reader, to really understand how a chess engine thinks. It will be quick, it could even be painless, so let us look at it in more detail. 

Down the very basics, a chess engine is composed by 3 fairly independent parts: the moves generator, the static evaluator and the search function.

The moves generator
First a remark: chess programmers often talk about chess positions as of FENs. These are just shorthands for describing things like "White has a king on the right bottom square h1, a pawn on the square c4 and a bishop on f2. Black only has a king on h8". Here is the FEN for the starting position
rnbqkbnr/pppppppp/8/8/8/8/PPPPPPPP/RNBQKBNR w KQkq - 0 1
The move generator gives you for any possible FEN all legal moves which can be done by the side to move.

The static evaluator
A very important element is to quantify how a position looks better than another one. This is done by the static evaluator. The static evaluator gives you for any FEN a value, representing its objective evaluation of the position. Traditionally, positive numbers mean white is good, negative numbers mean black is good. How could a static evaluator look like? It could be for instance something like: count all white pawns, subtract the number of black pawns and multiply by 10. In this case, for instance, both a position with 3 white pawns and 2 black pawns and a position with 4 white pawns and 3 black pawns will give you a value of +10. Obviously static evaluators of real chess engines are much more complex than that, having hundreds of elements. For instance Stockfish main evaluation function (Stockfish is the leading chess engine currently) has ca. 800 lines of C++ codes, with an element typically be coded in 5-10 lines, and this is only the main evaluation function. As I said, hundreds of elements.

The search
But the core of a chess engine is the search. Ahhh, the search. To understand how the search works, it is useful to make the following experiment. Imagine you couldn't play chess at all. Zero knowledge of the rules. The only thing you know is how to translate a chessboard position in a FEN. Plus, you are given a moves generator and a static evaluator. You start playing. You are white and you have to move. Since you have a move generator, you know what moves you could do, and you have a fantastic idea: you will just try them one by one and keep the one which gives you the best, in other words: maximal, static evaluation. If you want to spend some time, with this tool you can do the exercise with Stockfish. It turns out that moving the pawn in front of the king by two squares gives you a evaluation of 218. Cool! Let us move that pawn forwaaaaard!
But hang on! Black can do this too. Black would at its turn chooses the moves which maximize its value. This means, he will go through all possible replies and choose the one which minimizes my value. So it could be, in principle, that a possible reply to moving my king pawn by 2 squares give black a better value than all possible replies to moving my queen pawn by 2 values! 
In fact it is so, if I move my king pawn by 2 squares, the best he can do is to move his queen pawn by 2 squares and get a -20. In contrast, if I move the queen pawn by two squares, he can get is a +20 (again by moving its queen pawn by two squares). So, by minimizing on its turn, he invalidates my assumption that I just have to choose the move after which the static evaluation is best.
This is now bad, because it could be that there still another one of my possible 20 first moves that it is even better, in the sense that it makes the minimizing task for black even worse.
Now that we think about it, we should also consider that we also could reply to its replies, again trying to maximizing our value with our second move, so we have to also take this is consideration. And he could reply by minimizing and so forth and so on.
This loosely described algorithm is called minmax and is the basis of the chess engines. You just unroll all possible moves until, say, the 10th reply (technically: ply) you look at the values at the end of this unroll and both players choose alternately the minimizing or maximizing move. Minmax. Interestengliy, there is a variation which is equivalent and which is called alpha-beta which saves you some of the comparisons.

All competitive modern chess programs use some form of alpha-beta with variations and refinements. In this sense, they use brute force, in the very end. They just unroll the moves and do minmax. More brute force is not possible.

Now, if you do the math, you will notice that at every reply you can choose between 20-30 moves. This means, the possible positions after, say, 20 plies are 20^20, which is much more that you could store on your computer, even if every position and all information about it would stay on one bit. Alpha-beta gives you some more bang for the bucks, but not that much.

So, it was though that no computer will be ever able to reach the depth of the chess master. Oh, were they wrong...

domenica, maggio 06, 2018

What is a proof?

I was at lunch with some colleagues some days ago, and we started discussing about maths. Predictably, the conversation turned to the quite specific concept of proofs used by mathematicians. I had some insights, on the reasons for so many misunderstandings between mathematicians and others. I think many of them arise from a different understanding of the word "proof"

I want to exemplify on the classical result about the cardinality of the power set, so let us dive into math.

First, the power set ${P}(A)$ of a given finite set $A$ is the the enumeration of all the possible subsets of $A$.

If $A$ is, say, the set of $\{me, you\}$ the power set will contain:
  1. the empty subset $\emptyset$
  2. the subset containing only $me$
  3. the subset containing only $you$
  4. the subset containing both $me$ and $you$
It is a legitimate and somewhat interesting question: how many elements does the power set $A$ has, if $A$ has $n$ elements? Such counting questions are very common in math and constitute the basis of combinatorics. Let us look how we can come to a solution of the problem and, even cooler, a formal proof that the solution is correct.

Examples
The first step is, more often than not, to list some examples. For instance, if the set $A$ only contains one elements, the power set will contain:
  1. the empty set
  2. the subset consisting of that element
We now have an example for a 1 element set (2 subsets) and for a 2 element set (4 subsets). What about a 3 element set?

By the way, we have implicitly agreed on the working assumption that the cardinality of the power set only the depends on the cardinality of the original set. This is ok, if this is wrong, we will change it. (But yes, it will turn out that this is right).

Back to our example of our 3 element set $\{me, you, the\ others\}$. We have all which we have considered before:

  1. the empty subset $\emptyset$
  2. the subset containing only $me$
  3. the subset containing only $you$
  4. the subset containing both $me$ and $you$
and in addition, all of those plus $the\ others$

  1. the subset containing only $the\ others$
  2. the subset containing only $me, the\ others$
  3. the subset containing only $you, the\ others$
  4. the subset containing both $me, you$ and $the\ othes$
This totals to 8 elements. Let us summarize: for 1 element, 2 subsets, for 2 elements, 4 subsets and for 3 elements, 8 subsets. This looks suspiciously like $2^n$, with $n$ the cardinality of $A$. This will be our working hypothesis.

Insight
Having formulated our working hypothesis, we should ask ourself whether this can be possibly right. Upon thinking we recognize this: for every subset in our collection, every single element of $A$ can be present or not. Since we have $n$ elements, every element, we must make $n$ binary decisions of whether to include that element in the subset or not. This gives $2 \times 2 \times ... \times 2 $ possibilities. Yes! This is $2^n$. This is our insight, and typically it is enough if you are a physicists.
But it is not a proof. It is not a proof, as to be a proof it must be formalized in a standard way using the laws of logics. This is important, since this gives all mathematicians the chance of recognizing a proof as such.

The proof
Let us prove this, then. The standard way to prove stuff about natural numbers is the method of induction, which we will follow step by step. First we formulate our theorem in details.

If $A$ is a finite set with $n$ elements, then its power set $P(A)$ has $2^n$ elements.

We prove that this is true for 0. If the set $A$ is empty, then it has 0 elements, and one single subset, which is the empty set. $1  = 2^0$. Check.
To avoid being caught in misunderstandings and quarrels about 0 and empty sets, we will prove it also for a single element. This is not required, but we do it anyway. If $A$ has a single element, its subsets are the empty set and $A$ itself. $2 = 2^1$. Check.

Now let us go to the induction step. Let us assume, we known that if $A$ has $n$ elements, then we have $2^n$ distinct subsets of $A$. Let us reason about a hypothetical set $B$ with $n+1$ elements and let us pick one arbitrary element, which we will call $b$ out of $B$. Let us remove $b$ from $B$ and obtain $B'$ a set of obviously $n$ elements. We know from the induction assumption that $B'$ has $n$ elements and thus $2^n$ subsets, constituting the set $P(B')$. Let us now consider the set $X = \{S \cup \{b\} | S \in P(B')\}$.

Interestingly, $X$ and $P(B')$ are disjoint, as every set in $X$ contains $b$ which per construction absent from every set in $P(B')$. So, the cardinality of $X \cup P(B')$ will be the cardinality of $X$ plus the cardinality of $P(B')$.

It remains to show that
  1. the cardinality of $X$ is $2^n$
  2. $X \cup P(B') = P(B)$
For 1. we should give a bijective function for $P(B')$ to $X$, which turns out to be exactly the prescription with which we constructed $X$ in the first place. This kind of clerical tasks is typically left to the reader as an exercise, and we will comply with the mathematical tradition.

For 2. we notice that if $Y$ is a subset of $B$ it either contains $b$ or not. If not, it is in $P(B')$ and thus in  $X \cup P(B')$. If yes, it equals to $Y \setminus b \cup \{b\}$. Now, $Y \setminus b$ only consists of elements of $B$ apart of $b$, so it is in $P(B')$ and when we unit sets in $P(B')$ with $\{b\}$ we get by construction sets in $X$, so $Y$ is a set in $X$ and thus in $X \cup P(B')$. We just proved that $P(B) \subset X \cup P(B')$. The converse direction is left to the reader an exercise, as it follows the same principles. We also do not want to spoil him of the incredible feeling of finishing a proof.

Conclusion
What I hope I could convey is that sheer amount of work needed to go from an insight, giving us just a working hypothesis and some ideas, to a complete mathematical proof. The real question could be here: why bother? Indeed our insight was right, and the loosely described proof ($n$ binary decisions giving $2^n$ different possibilities) is perfectly fine.

I think the problem is often that this kind of insight often fails with more complex arguments, even in seemingly intuitive cases.

Formalized, standardized proofs in maths are a powerful method to fight the natural tendency of the human mind to fall victim of cognitive illusions and are thus one of the milestones of the human thinking!


mercoledì, maggio 02, 2018

White King, Red Queen: selling the chess soul, part II.

We just left 1979, with D. Hofstadter being skeptical and not at all enthusiastic of a chess computer. Unimpressed by the opinions of the American scientist, in the German city of Munich, exactly at the same time, a small company just hired two programmers to bring to the market the first portable chess computer.

Chess was quite popular at that time. The famous Spassy-Fischer match of 1972 had been an incredible world-wide drama. Subsequently,  Fischer started his path into madness, was deposed and the rivalry between the two soviet giants Karpov and Kasparov began.

So, a short time after the introduction of the first personal computers, in the middle of a world-wide chess pandemic, you could by a small box of plastic, which was able to play chess at the level of a mediocre, but loyal and always available club player. It was called Mephisto I.
Mephisto I was a real blockbuster. You could buy for 500 DM (or 500 € of today), well below the price of a personal computer back then. It would help you going through tactical variations in post-morten analyses of games during chess-club evenings. They would sell 100.000 pieces per year only in Germany. The company producing the Mephisto, Hegener + Glaser AG, enjoined 10 years of almost complete monopoly of the chess computer market, and the financial reward coming from it.

And as it often happens, they were killed by their own inertia. The market was changing: more and more chess computers were available, the once esoteric mystery of alpha-beta search well known, and most important, the 486 entered the scene. A 486 with a decent chess program could reach chess levels worth of a national master. You could'nt bring it to the club, right. But for correspondence player it was a cheap insurance against tactical blunders. Chess programs entered the life of serious chess players and their moral authority dictated that Mephisto and all portable computers should be abandoned by the weaker chess players, or be condemned not to learn chess in their entire life. So they abandoned.

Why did'nt Hegener + Glaser AG see this coming? They had the professional chess programmers, they should have noticed at latest in 1989, as no chess computers entered the world computer chess championship or in 1991 when Mephisto lost the title to a software version of itself. Nobody knows.

H + G survived until 1994, and here we leave them and their 28 million DM of debts as a further victim of the endless race ordered by her majesty the Red Queen.


domenica, aprile 29, 2018

White King, Red Queen: the arms race in computer chess, part I.


Some weeks ago I went through a very insightful book about the red queen hypothesis: the evolutionary arms race between coexisting species, or between different gender in the same species, or finally, between individual in the same gender competing for offspring.

One classical example of the first one is the coevolution between figs and wasps: each fig having more or less a single dedicated wasp able to fertilize it. Those two coexisting species evolved through mutual competition: like inhabitants of the wonderland, they cannot stand still and are forced to continuously evolve in a perpetual arms race.

I want to now make the bold claim that the red queen hypothesis does not hold true only in the context of biological evolution, but in all cases in which there is competition and possibility of change. I would like to exemplify this bold claim on the amazing story of chess engines.

This is a community known the mosts only through the fame of the mythological Deep Blue, the first chess computer to defeat a human World Champion, the acclaimed Gary Kasparov.

I make this choice as it happens that I am active in the chess programming community since some years. Having been a mediocre club chess player in my youth, I have always been fascinated with chess engines and artificial intelligence. This even determined some of my career choices.

This will be a weird reading for people that never heard of chess tournaments or chess programming. You will get to know a competitive environment pushed forward by both by financial interests and the desire for intellectual prestige.

As a bonus, you are going to learn something about one the longest lived and most united communities in the digital world, the one of chess enthusiasts and programmers. To quote the motto of the World Chess Federation: Gens Una Sumus. We are one people.

I would like to start my journey with Gödel, Escher, Bach, an eternal Golder Braid, the gorgeous book by D. Hofstadter about everything, including artificial intelligence. There is one place in the middle of the book, where Hofstadter, a leading cognitive scientists of that time, starts to write about chess and chess computers. He is skeptical about the possibility that a computer program could defeat a human chess master (reddit link), adding that this is probably due to the peculiar way in which chess masters think. This was 1979. The AI scientists of that time seemed to predict a long hard time for chess computers. Obviously, if one is to negate such a development, this means that the development is already in course. There is a reward, financial gain and intellectual prestige, and there is a possibility to change, the computer technology rapidly evolving in those years.

Enter The Red Queen.