Showing posts with label Engines. Show all posts
Showing posts with label Engines. Show all posts

26 August 2023

Chess018 and the CCRL

In last week's post, Chess018 Is a Thing (August 2023), I suggested an idea for a follow-up post:-
One more idea that could be easily tested is to look at the chess018 start positions (SPs) on CCRL. Which SP is statistically best? Which is worst? What are the most popular first moves for each SP? Sounds like another post is taking form.

The following chart lists the 18 SPs in the chess018 family. The first three columns are from my own database of SPs. The last three columns are from the CCRL (see the sidebar for a link).

The column 'Dist' ('Distance') is a metric I whipped up when I first started investigating chess960. It was explained in the following posts:-

Some day I hope to find a use for 'Dist', e.g. a correlation with some other observation. Today is not that day.

The CCRL numbers for chess018 should be compared to the same numbers across all 960 SPs. At first glance, nothing unusual appears. Somewhat curious is the one SP where the score for White is less than 50% -- the traditional start position. I imagine it has something to do with the engines not using a book.

22 July 2023

TCEC FRD

A year ago on this blog I had a pair of related posts...

...where 'FRC' stands for 'Fischer Random Chess'. Last month on my main blog I had a pair of follow-up posts...

...'FRC' is, of course, another name for chess960, where the start positions for the two players are mirrored. 'DFRC' is 'Double FRC', where the two players have different start positions.

In that post 'Stockfish Wins TCEC DFRC2', I noted that future TCEC FRC/DFRC events will use a format they call 'FRD', meaning 'Fischer Random Double'. For an explanation of FRD, see that post. Don't be surprised if this current post is the last mention of FRD on this blog.

24 June 2023

Chess960 in the Cloud

In a recent post, When Chess960 Reduces to Chess (May 2023), I wrote,
The point where castling is no longer an option is exactly where chess960 starts to look and feel like chess starting from the traditional position (SP518 RNBQKBNR). This is the point where a chess service like Chessify becomes fully viable.

A few weeks before writing that post, I had the opportunity to use Chessify on live chess960 positions. I was traveling for two weeks without any of the software tools I normally use for correspondence chess960. Before leaving home, I loaded six games into Chessify by using their 'Import' function to enter the current FEN position for each of the games. Most of the games had already reached the point where castling was no longer an option, but there were a couple of exceptions. A week into my trip I finally had some time to look at the games.

In the diagram's top position White has lost the castling privilege, while Black can still castle to both sides. I was playing Black. Stockfish NNUE on Chessify told me that 14...g5 was by far the best move, leading to a slight advantage for White. I vaguely remembered from my previous analysis that Black had an advantage in the position and decided that it must have something to do with castling. Since ...O-O-O isn't happening any time soon, the correct line must involve ...O-O. The move 14...O-O is a blunder because of 15.Nd7, which suggests 14...Nb6 to protect d7. Using Stockfish to continue the analysis and manipulating the FEN to simulate ...O-O, I decided that 14...Nb6 must be the right move and played it.

When I returned home, I looked at my previous analysis and discovered that 14...g5 was indeed the right move, leading to a significant advantage for Black. With 14...Nb6, I had spoiled my position and gone from a probable win to a struggle for a draw.

In the bottom position White has just castled, while Black can still castle ...O-O-O. I was again playing Black. Still using Stockfish NNUE, I decided that 15...O-O-O was premature and played 15...f6 instead. Back home I was happy with this move and after 16.f4, continued 16...O-O-O.

A few moves later I started to see weirdness from Chessify's Stockfish. I always calculate multiple variations (MPV in chess jargon) and reached a position where around depth 25, a key move disappeared from White's candidate moves and another move appeared twice. With the help of Chessify support (top notch in my opinion) I discovered that the FEN for the position was wrong. It indicated that both White and Black could still castle to the side where they had already castled. I corrected the FEN and the Stockfish analysis appeared normal. I decided that there was a bug in Stockfish triggered by faulty FEN.

In my four other games nothing unusual happened. While traveling I made a move or two in each game and when I returned home confirmed that my play was correct. My conclusions from this exercise? 1) Playing using only cloud tools is feasible; and 2) Playing chess960 with an engine that doesn't understand chess960 castling is possible, but risky. Always double check the chess960 FEN...

20 May 2023

When Chess960 Reduces to Chess

Let's talk turkey. Once in a while I like to document my recent experience playing chess960 online. Earlier this year I posted The Fascinating World of Chess960 (January 2023), where I wrote,
Last month's post was also about switching to a different online service for playing chess960. [...] I continued playing on LSS until last year. I was playing chess960 in a couple of multi-stage events, where success in one stage promotes a player to the next stage. I decided to skip the next stages, essentially taking a year off from serious play. [...] I switched to Chess.com in May 2022, playing one or (maximum) two games of correspondence chess at a time.

A month later I wrote a post about a Chess.com service, Chess.com Reviews a Chess960 Opening (February 2023). Since then I've stopped playing Chess.com for reasons that I won't discuss in this post. I went back to LSS, partly with the intention of evaluating Chessify for chess960.

This is the first time I've mentioned Chessify on this chess960 blog, although I've discussed the service several times on my main blog. In Chessify Resources (March 2023), I wrote,

The main problem with chess960 in a traditional chess environment stems from the castling rules. Since chess960 games tend to become extremely tactical after a few moves have been played, there is nevertheless some value in trying to confirm the tactics with a traditional, non-chess960 engine. [...] I'll continue using Chessify to look at chess960 positions.

There are three phases of a chess960 opening (often overlapping with the early middle game):-

  • Both sides can castle.
  • One side loses the castling privilege.
  • The other side loses the castling privilege.

I say 'loses the castling privilege', because it can arise when castling, when the King moves, or when both Rooks move. The point where castling is no longer an option is exactly where chess960 starts to look and feel like chess starting from the traditional position (SP518 RNBQKBNR). This is the point where a chess service like Chessify becomes fully viable.

22 April 2023

Leela in TCEC FRC Events

My previous post Breaking the 4000 Barrier (April 2023) was about chess960 ratings for engines. I wrote,
The fourth [rating] list was based on chess960 games. Here are the top 25 engines from that list.

I received a question -- @nicbentulan: Is Leela on this list? (twitter.com) -- against the Twitter anchor for the post. A quick look at the list reveals that Leela is on the most recent list at no.42, rated 3008:-

no.42 • Lc0 0.29.0 CPU_744706 • 3008

That rating is nearly 1000 points lower than no.1 'Stockfish 15.1', currently rated 4005. The Lc0 rating looks dubious, given that Leela is currently one of the top three engines in the world. How has it done in TCEC FRC competitions? Following is a list of posts from this blog:-

  • 2022-07-30: TCEC C960 FRC5 • In the 'Final League', Stockfish and LCZero finished in a tie for 1st and 2nd to qualify for the 'Final Match'. [...] In the 'Final Match' Stockfish beat LCZero +17-13=20.'
  • 2022-01-22: TCEC C960 FRC4 • 'Stockfish and LCZero tied for 1st/2nd in the [TCEC] FRC4 'Final League', a point ahead of KomodoDragon. Stockfish beat LCZero +13-9=28 in the Final.'
  • 2021-03-27: TCEC C960 FRC3 • 'In the 'FRC 3' final, KomodoDragon beat Stockfish by a score of +2-1=47.'
  • 2020-12-26: TCEC FRC2 • 'In the FRC2 Final League, LCZero and Stockfish finished first and second to qualify for the 50-game final match. Stockfish beat LCZero +8-0=42.' [...] 'The first of the three posts above linked to Stockfish, the Strong (July 2014) on this blog, plus two other followup posts based on FRC1. FRC1 was held three and a half years before AlphaZero made waves with its revolutionary AI/NN technology, soon to be followed by Leela Chess Zero (aka LCZero / LC0).'
  • 2014-08-02: TCEC Season 6 - Chess960 • 'After posting Stockfish, the Strong, winner of this year's TCEC Season 6 Special Event (chess960), I started looking at the games from the event. I couldn't find a crosstable, so I made one myself, shown below.'

Why is Leela so far down on the CCRL rating list? I'll leave that for the CCRL to answer.

15 April 2023

Breaking the 4000 Barrier

The title comes from a recent post on my main blog Breaking the 3600 Barrier (April 2023), where the '3600 Barrier' refers to a chess rating level. In that post I wrote,
I see that there are four CCRL rating lists. Shown below are the top five engines from three of the lists.

Why didn't I show all four lists? Because the fourth list was based on chess960 games. Here are the top 25 engines from that list.


CCRL 40/2 FRC [C960] - Index

The most striking feature of the list is that the top engine, Stockfish 15.1, is rated over 4000. This is more than 250 points higher than the top rating shown on the '3600 Barrier' chart. In fact the top four engines on the CCRL FRC list are higher then the top engine on the earlier list. I couldn't determine why this is.

[NB: The domain given here (ccrl.chessdom.com) is not the same as the domain in the right sidebar (computerchess.org.uk), although the two sites appear to be identical. The ratings are based on the same database of games used to compute the CCRL 'Opening statistics'.]

25 February 2023

Chess.com Reviews a Chess960 Opening

In last week's post, Chess.com Pinpoints a Tactical Error (February 2023; see the post for a copy of the game's PGN), I used 'Chess.com's Game Review Tools' to find out where I had made the first mistake in losing a chess960 game. This week I'll use the same tools to extract comments on the opening moves.

The following diagram shows the start of the same game featured in the 'Tactical Error' post. There's more I can say about the look and feel of the 'Analysis' tool itself, but I'll save that for a series I'm doing on my main blog. The most recent post in that series was Chess.com's Game Review Tools PGN (February 2023).

Shown on the left is the start position for the game, 'SP350 NRKQRBBN'. On the right is a summary of the overall quality of the players' moves ('Brilliant', 'Great Move', ..., 'Blunder'). For example, the tool considers that both players made one 'Great Move'.


AV vs. bemweeks | Analysis (chess.com)

The following table shows the tool's comments on the first 12 moves of the game (24 ply deep). I stopped the analysis when I reached the move where I committed the 'Tactical Error'.

Move Short
Comment
Long
Comment
Eval  
1.e4 is excellent This prepares the bishop for development. This threatens to reveal an attack on a pawn. +0.13
1...e5 is good This prepares your bishop for development. +0.30
2.f4 is excellent This exposes an attack, threatening a pawn. +0.27
2...Nb6 is excellent Your piece jumps in to protect a pawn! +0.41
3.fxe5 is best Right on target. +0.41
3...f6 is an inaccuracy You are threatening to attack a trapped rook. +1.09
4.Nb3 is a mistake This loses a pawn. +0.05
4...fxe5 is best That wins a free pawn! +0.05
5.Ng3 is excellent One of the best moves. -0.02
5...Qf6 is best You activate your queen by moving it off of its starting square. -0.02
6.Be3 is good This moves the bishop to a better location, allowing it to control more squares. -0.30
6...O-O-O is good Your rooks can see each other now, allowing them to provide mutual defense. +0.02
7.d3 is best That's what I would have recommended. +0.02
7...d5 is best You are threatening to kick a bishop. +0.02
8.Qd2 is an inaccuracy This ignores a better way to develop a queen off its starting square. -0.47
8...Qc6 is good You are threatening to kick a bishop. -0.10
9.Bg5 is good This wins a tempo by threatening a rook and forcing it to move away. -0.38
9...Be7 is a mistake You are threatening to win material. +0.56
10.Bxe7 is best After all captures, this is an equal trade. +0.56
10...Rxe7 is best You trade off equal material. +0.56
11.Nf5 is good This attacks a rook, winning a tempo when it moves away. This threatens to fork pieces. +0.27
11...Red7 is good You have now doubled your rooks, allowing them to team up to create threats. +0.62
12.exd5 is best This exposes an attack, threatening a pawn. +0.62
12...Qxd5 is a mistake You overlooked a better way to recapture a piece. +1.63

There's much more I could say about the comments, but it would not be useful at this point. Here are a few comments that jumped off the screen at me.

  • 8...Qc6; 'You are threatening to kick a Bishop.' • The move defends against a nasty x-ray threat.

  • 9.Bg5; 'This wins a tempo by threatening a Rook and forcing it to move away.' • The move doesn't win a tempo, but it might lose a tempo by forcing the Rook to a better square.

  • 9...Be7; 'Is a mistake. You are threatening to win material.' • It's a mistake to win material? Something does not compute here.

And so on. The long comments are generally lame and show little understanding of chess960 opening objectives. What happened in the fight for the center?

The most valuable part of the exercise is to see the change in evaluation from one move to the next. It reaffirms the severity of my mistake on the 12th move.

18 February 2023

Chess.com Pinpoints a Tactical Error

In last month's post, The Fascinating World of Chess960 (January 2023), I discussed a shift in focus for my own chess960 games. In a nutshell, to play chess960 I switched from a site allowing engines to a site forbidding them. I finished the post saying,
That's the background for a series of posts that I plan to write for my games on Chess.com. There are several aspects to be covered, e.g. Game review tools

I introduced those tools on my main blog in Chess.com's Game Review Tools (February 2023), using an example chess960 game. I'll use the same game in this current post. After four wins, it was the first game I lost on Chess.com starting June 2022, when I adopted the no-engine approach.

The 'Game Review Tools' post used numbers to identify the different screen shots and I'll follow the same numbering scheme in this current post. There are three different review tools that I numbered '02', '03a', and '05a'. The '02' tool shows the moves of the game and the times used for each move. It also provides (1) an entry into the '03a' and '03b' tools, and (2) a PGN download of the moves of the game, without variations or comments.

At this point, I don't see much difference between the '03a' and '05a' tools. The differences seem mainly cosmetic, so I'll continue with the '05a' tool. It's accessed via a feature called 'Saved Analysis'. In the game I was outplayed tactically and didn't know where I had gone wrong.

I had Black in a game starting 'SP350 NRKQRBBN'. I was pleased with my position after the first few moves and thought that I had equalized, maybe even gained a slight advantage. Then suddenly I had an inferior game. Why?

The '03a' and '05a' tools offer commentary on each move played in the game. I won't discuss the early comments in this post, because I'm not yet convinced they are helpful. The critical position is shown in the following screenshot, Black to move.


AV vs. bemweeks | Analysis - Chess.com • After 12.e4-d5(xP)

White has just captured a Pawn on d5 with 12.exd5, and Black has four possible recaptures. The move 12...Rxd5 is a '??' blunder, allowing the family fork 13.Ne7+. That leaves three other moves. Since White's last move also discovered an attack on the e5-Pawn, I played 12...Qxd5, protecting that Pawn. That was a mistake, which the tool flags with the remark '(?) Qxd5 is a mistake'.

What's better? The tool suggests 12...Nxd5. Now if 13.Rxe5, Black has 13...Qf6, when White is in trouble. I didn't see that possibility during the game and never recovered. Following is the PGN as provided by the '02' tool.

[Event "Let\\'s Play! - Chess960"]
[Site "Chess.com"]
[Date "2022.07.05"]
[Round "?"]
[White "Andreasvinckier"]
[Black "bemweeks"]
[Result "1-0"]
[Variant "Chess960"]
[SetUp "1"]
[FEN "nrkqrbbn/pppppppp/8/8/8/8/PPPPPPPP/NRKQRBBN w EBeb - 0 1"]
[WhiteElo "1977"]
[BlackElo "1907"]
[TimeControl "1/86400"]
[EndDate "2022.07.23"]
[Termination "Andreasvinckier won by resignation"]
[initialSetup "nrkqrbbn/pppppppp/8/8/8/8/PPPPPPPP/NRKQRBBN w EBeb - 0 1"]

1. e4 e5 2. f4 Nb6 3. fxe5 f6 4. Nb3 fxe5 5. Ng3 Qf6 6. Be3 O-O-O 7. d3 d5 8. Qd2 Qc6 9. Bg5 Be7 10. Bxe7 Rxe7 11. Nf5 Red7 12. exd5 Qxd5 13. Ne3 Qc6 14. g3 Be6 15. Qa5 Bxb3 16. axb3 e4 17. Bh3 exd3 18. O-O-O dxc2 19. Bxd7+ Rxd7 20. Rxd7 Qxd7 21. Qxa7 Nf7 22. Qa5 Nd6 23. Nd5 1-0

The Chess.com tools offer several different PGN downloads. I'll discuss those in another post on my main blog.

21 January 2023

The Fascinating World of Chess960

Last month's post, Christmas Eve (December 2022), wasn't just about creating a new category for Posts with label MW's games (that's me). It was also about switching to a different online service for playing chess960.

I recorded my first game of chess960 in a post on my main blog, titled Chess960? I'm Hooked! (September 2008). The PGN embedded there says the game was played on SchemingMind.com (SM), a site for correspondence chess that is particularly strong in its support of chess variants. I continued playing there until 2016.

A few years after getting hooked, I started playing on another correspondence site, The Lechenicher SchachServer (December 2012; LSS), for reasons explained in that post. I continued playing on both SM and LSS until 2016, when I ran into a problem on SM and decided to leave. To make a long story short, SM has a no-engine policy, but makes little effort to enforce it. LSS allows engines for most of its events and I preferred the clarity of LSS.

I continued playing on LSS until last year. I was playing chess960 in a couple of multi-stage events, where success in one stage promotes a player to the next stage. As the two events were winding down, the site announced the next stages. Unfortunately for me, the start of both events coincided with a pair of two-week vacations that I had been planning for some time. Since LSS events allow only two weeks of vacation on a fixed number of days for a game (no increments), I was faced with an immediate time deficit in all new games. I decided to skip the next stages, essentially taking a year off from serious play.

As my active games gradually came to a conclusion, in my free time I started using an engine to analyze my old correspondence games from the pre-engine, pre-chess960 era. I was amazed that my moves were generally approved by modern engines. I could often recollect the reasoning and emotions behind my moves and realized that using an engine had turned me from a chess player into an engine operator. After 14 years of playing chess960, I hadn't gained much insight into its subtleties, because I was essentially playing what the engine instructed me to play, often without understanding why.

I decided to switch to a site that didn't allow engine use. SM was out because it doesn't enforce its policy. Then I remembered Chess.com, which has a good reputation for vigorously enforcing its no-engine policy, even if it leads to controversial decisions. I had played a few games of chess960 there in 2009-2010 and more recently in 2019, an experience documented in Playing the FWFRCC (June 2019).

I switched to Chess.com in May 2022, playing one or (maximum) two games of correspondence chess at a time. What a difference! Where my last years with LSS involved struggling against players with far more powerful engines than I was using, at Chess.com I was using my own head to play real chess against other players doing the same. After all, that's what had attracted me when I first started playing chess so many years ago.

So far I've played about a dozen games on Chess.com, never once tempted to use an engine. I also know full well that if I do use an engine and am caught, I will lose the premium membership that CEO Erik gave me when I was writing a review of the site for About.com in mid-2008.

That's the background for a series of posts that I plan to write for my games on Chess.com. There are several aspects to be covered:-

  • Insights from my games
  • The correspondence play interface
  • Game review tools
  • The site's custom anti-cheating measures
  • And more...?

I'll wander through these topics in future posts, some of them on my main blog. Thanks to both SchemingMind and LSS for the terrific support of chess960 throughout the years. May they continue to introduce keen chess players to the fascinating world of chess960.

27 August 2022

TCEC DFRC1

In a perfect world, this post would be a followup to the previous post 2022 FWFRCC Kickoff (August 2022; 'FIDE World Fischer Random Chess Championship'). That event will undoubtedly dominate the blog for the next few months, but first I have to tackle an outstanding topic before it becomes old news.

This post is more of a followup to last month's TCEC C960 FRC5 (July 2022), because it's also about the TCEC. It's in reference to a post on my main blog, Stockfish Wins TCEC DFRC1, Leads CCC18 Rapid Final (August 2022). As a summary of DFRC1, I wrote,

In DFRC1 ('Double Fischer Random Chess: (960*960) possible starting positions'), Stockfish and LCZero finished tied for 1st/2nd places with 16.0/22, 1.5 points ahead of KomodoDragon, which was 2.0 points ahead of Stoofvlees. In the 50 game final match, Stockfish beat LCZero 29.5-20.5 (+18-9=23). For more discussion of the event, see the next post on my chess960 blog.

I know, this isn't the 'next post' on this blog; it's the 'next+1 post'. The unexpected news about the FWFRCC took priority. To understand what the TCEC accomplished -- and it's without question a noteworthy accomplishment -- let's quote some TCEC !definitions [Warning! : Serious chess engine jargon ahead...]:-

!dfrc • Double Fischer Random Chess: The same as Fischer Random Chess [FRC; C960], except the White and Black starting positions do not necessarily mirror each other. Double FRC has 921,600 (960*960) possible starting positions. [Also has links to files for DFRC 'Openings' and 'Evals'.]

!dfrctest • Initial test of !DFRC at TCEC. To check a) engines handling of DFRC - see !bugs, b) some DFRC start positions, c) cutechess handling of DFRC, d) GUI handling of DFRC. Time control: 7 minutes + 1 second/move. Duration 1 week

!dfrc1 • 1.SF (on SB 196) 2.Lc0 (SB 192.5) 3. KD. 26 participants (FRC5 + Bagatur), Swiss format, games from starting positions, 11 rounds, and 30min+3s TC. Estimated duration: 12 days.

!bookdfrc • Selected 1600 opening positions for the DFRC event, and for the DFRC testing. These are the positions for which KD, Eth, Berserk and SF (the latter through dbcn) agree that white has about a 50% chance for a win. The selection was then further pruned using careful SF analysis.

!unbust • Took positions where starting position abs(eval) >= 2.00, 17739 DFRC positions. Played Stockfish 202206020749 200k vs Stockfish 100k and 8071 were still 'busted'. Played those 1M vs 100k and 1172 were left 'busted'. Played those 10M vs 100k and 83 were left 'busted'. Finally, those 100M vs 100k and 10 are still 'busted'. These will be used in Bagatur vs Stockfish bonus.

!chess324 • A subset of DFRC where the Kings and the Rooks are at the usual starting position. Since castling is standard, this allows all engines to play.

For more about that last definition, see Chess324 (talkchess.com; lkaufman, aka Larry Kaufman of Komodo++ fame). The dean of chess engines explained, 'The Kings and Rooks are placed on their normal positions. All the other pieces for White and Black are placed randomly, with no symmetry requirement, the only restriction being that for each side the Bishops must be on opposite colored squares.'

Chessbase weighed in on a related topic with Double Shuffle Chess: a fun variant against Fritz Online (chessbase.com; Albert Silver). In this variant the Kings aren't required to be between the Rooks and there is no castling. No thanks, I'll pass. Maybe the TCEC or the CCC will give it a try.

After that little detour, let's get back to the TCEC. For archive info (crosstables, PGN, etc.) about the DFRC1 events, see:-

Since I'm already following the TCEC, I'll keep an eye on their DFRC events. My first impression is that it's more for chess engine enthusiasts than for chess960 enthusiasts, but I was no fan of chess960 when it first appeared on my radar either.

30 July 2022

TCEC C960 FRC5

Earlier this month on my main blog, I posted Stockfish Wins TCEC Swiss 3 and CCC17 Blitz (July 2022). After 'TCEC Swiss 3':-
The site then launched FRC5, which is currently in the four engine 'Final League' stage (KomodoDragon is missing).

The 'Final League' served as a qualification event for the two engine 'Final Match'. The following chart shows the crosstables for both of the 'Final' events, which have completed.


Top: Final Match
Bottom: Final League

In the 'Final League', Stockfish and LCZero finished in a tie for 1st and 2nd to qualify for the 'Final Match'. That result came about after they tied their individual mini-match +4-4=0 and crushed the two last-placed engines in their other mini-matches. In the 'Final Match' Stockfish beat LCZero +17-13=20.

As for the remark, 'KomodoDragon is missing', the other member of the current engine triumvirate was eliminated at the semifinal stage. Through a specific combination of first and second place finishes in the groups of the preliminary stage, all three engines played in the same four-engine semifinal group. LCZero finished a half point ahead of Stockfish, which was a full point ahead of KomodoDragon, thereby eliminating KomodoDragon. For more about the specifics of the competition, see:-

An important section of the rules is worth highlighting:-

4. Openings books
a. All matchups until the Semileagues stage will be played from a randomly generated FRC aka Chess960 start position.
b. For the Semileagues, the Final League and the Final a shallow book will be used.

What does 'shallow book' mean? I couldn't find an explanation on the TCEC Wiki, but the site's !commands inform,

!mob • MOB - Minimalistic Openings Book made by Kan. The idea is to leave lots of choice to the engines and be as shallow as possible while providing an optimum of variety.

I imagine that definition was developed for the traditional start position SP518 and has been reused for chess960. My post on the previous event, TCEC C960 FRC4 (January 2022), looked at TCEC practices for chess960 opening books. To adapt a phrase from that post,

The researcher behind the 'shallow' analysis should make available his full analysis showing which positions were eliminated for which reasons.

That FRC4 post also referenced a couple of my own 'Iceberg' posts (November 2021) analyzing engine runtime data from TCEC FRC3. Add FRC5 to the backlog of events for this sort of analysis.

22 January 2022

TCEC C960 FRC4

In a recent post on my main blog, Stockfish Wins Both TCEC FRC4 and CCC16 Bullet Events (January 2022), I made three observations related to chess960:-
(1) 'Stockfish and LCZero tied for 1st/2nd in the [TCEC] FRC4 'Final League', a point ahead of KomodoDragon. Stockfish beat LCZero +13-9=28 in the Final.'

(2) 'A [TCEC] note mentioned, "!bookfrc • Final League and the Final will use unbalanced books [...] On the edge between draw and white win." For more info, see TCEC FRC 4, under 'FRC Book Generation'.'

(3) 'After FRC4, the site ran an event called 'S22 - DFRC Sanity Check'. What's DFRC? "!dfrc • Double Fischer random chess: The same as Fischer random chess, except the White and Black starting positions do not mirror each other. Double FRC has 921,600 (960*960) possible starting positions."'

The 'more info' reference in (2) was for TCEC FRC 4 - TCEC wiki (wiki.chessdom.org), where the nuts and bolts of the tournament are explained. I covered the previous event in TCEC C960 FRC3 (March 2021).

One welcome difference between FRC3 and FRC4 was the increased number of competitors in the 'First phase', comprised of four leagues. In FRC3, there were four engines in each league; in FRC4, six engines were planned, although only 23 engines started the event. After FRC3, I analyzed engine runtime data for the first time in a pair of posts:-

  • An Engine Iceberg (November 2021) • 'TCEC FRC3 was a 50 game match won by KomodoDragon over Stockfish on a final score of +2-1=47.'
  • The Engine Iceberg Looms Larger (ditto) • 'Final match of the CCC C960 Blitz Championship (October 2021).'

It might be useful to repeat the exercise for FRC4, although I should be clear on objectives for the exercise. What can be learned by looking at only a small, random subset of the 960 possible start positions?

The TCEC wiki page discusses 'unbalanced books'. It's an interesting concept, but perhaps too heavy on the human manipulation: 'following work is done by hand'; 'then (by hand and eye) I choose'; '[sequences] that don't look crazy to me'; 'eliminated lines that looked too drawish or too busted'; 'some looked too artificial, some looked a bit too similar to others'.

Traditional A/B engines have never been particularly good at evaluating the long term consequences of opening decisions. A comprehensive analysis extends well beyond their search horizons. Maybe the AI NNUE engines are better at this, but that hasn't been studied anywhere (that I know of).

A red flag goes up when I see a phrase like 'lines that looked ... too busted'. In the years of writing about and playing chess960, I haven't seen any start positions that were 'too busted'. To the contrary, Black always has resources to counteract White's various initiatives. Perhaps the researcher behind the analysis (Bastiaan) should make available his full analysis showing which positions were eliminated for which reasons.

Another phrase caught my attention: 'not a single position favours Black in my analysis'. This is what one would expect to see in a position between opponents having exactly the same resources, except one gets to move first. Otherwise we would talk about 'first move disadvantage' or 'zugzwang in the start position'.

As for DFRC (FRC squared? chess921.6K?), this is a new area for analysis. It's another example of Gene Milener's idea that I covered in Chess960 Phase Zero (November 2018). A first action might be to examine runtime data from the DFRC games.

27 November 2021

The Engine Iceberg Looms Larger

Last week's post, An Engine Iceberg (November 2021), looked at runtime data from 'TCEC C960 FRC3', which I covered on this blog in March 2021. I wrote,
Since the data covers only the first move of 25 SPs (50 games) out of the full set of 960 SPs, it's obviously just scratching the surface. Suppose we had data for the first few moves of all 960 SPs from many different engines played over a long period of time. What might we learn from this?

For this current post I repeated the exercise on the final match of the CCC C960 Blitz Championship (October 2021). I wrote,

In the final match Stockfish beat Dragon +10-1=589. Yes, more than 98% of the final games were drawn.

I loaded the PGN for all 600 games into my database and ran a preliminary analysis. There were two small surprises.

The first surprise was that the data for individual moves was not the same for both the TCEC and the CCC. Here are examples for the first move of the first game in both events.

TCEC: 1. e4 {d=36, sd=36, mt=147236, tl=1657764, s=81363821, n=11979602252, pv=e4 Nb6 Nb3 e5 g3 g6 Ne3 c6 f4 exf4 gxf4 f5 exf5 gxf5 Bf2 Qf6 c3 Nd5 Bc2 Nxe3 Qg1 Ne6 Bxe3 Bc7 O-O-O O-O-O Nd4 Bf7 Rf1 Bh5 Rde1 Nxd4 Bxd4 Qf7 b3 Be2 Rf2 Bg4 Kb2 Rxe1 Qxe1 Rg8 Qf1 Bb6 Bxb6, tb=0, h=99.9, ph=0.0, wv=0.26, R50=50, Rd=-11, Rr=-1000, mb=+0+0+0+0+0,}

CCC: 1. d4 {+0.45/32 9.6s, ev=0.45, d=32, pd=g6, mt=00:00:09, tl=00:04:55, s=148396 kN/s, n=1415409674, pv=d4 g6 e3 d5 g4 c6 c4 dxc4 f4 g5 fxg5 Na6 Nf2 e5 Bxc4 Nb4 Na3 exd4 Qf3 Be7 exd4 Qxd4 O-O O-O Bb3 Ne6 h4 Qg7 Be3 Nd5 Ne4 Nxe3 Qxe3 h6 Nc4 hxg5 Ncd6, tb=0, R50=50, wv=0.45}

Fortunately, the important 'wv' and 'pv' fields are available for both events. Any other fields I decide to use might require some sort of conversion.

The second surprise was that the CCC start positions (SPs) were not repeated for a second game, colors switched, between the engines. Instead, a new SP was assigned to each game. The left table in the following chart shows that some SPs were nevertheless repeated up to five times.

In addition to the six SPs shown in the table, 24 SPs were repeated three times and 90 were repeated twice. I assume that the SPs are chosen randomly for both the TCEC and the CCC, perhaps with the exception of SP518 RNBQKBNR, but I know from past investigations that several bad algorithms are in use elsewhere; see Start by Placing the Bishops (September 2017) for examples.

The center table in the chart shows the number of times a certain first move was chosen across all 600 games. For example, the initial moves 1.a4 and 1.b3 were both chosen 19 times. Just as in SP518, advancing a center Pawn two squares (1.c4, 1.d4, ...) is the most popular opening strategy. Although any single SP has a maximum of four initial Knight moves, sometimes only two or three moves, all eight moves are possible across the 960 SPs.

There are a number of questions for further exploration. When is the advance of an edge Pawn -- 19 x 1.a4 or 5 x 1.h4 -- desirable? I suspect these are position where the Queen starts in the corner behind the Pawn. Why the large difference between the counts on the two edge Pawns? Perhaps this is because of castling O-O/O-O-O considerations. Also worth noting is that O-O/O-O-O was never chosen for the first move.

The rightmost table in the chart gives a rough distribution of initial 'wv' values, i.e. what value did the engine calculate for its first move? These are truncated values, e.g. the CCC 'wv=0.45' shown above is counted in the table as 'wv=0.4'. I could have used roundoff and a bar chart to display the counts more accurately, but I ran out of time.

One big question presents itself here. Why are there so many 'wv' greater than 0.5, but so few decisive results during the match? I also need to determine if Stockfish and Dragon calculate values in the same statistical range. I doubt that they do.

The three tables in that chart lead to many questions and few answers. I'll take this up again some other time.

20 November 2021

An Engine Iceberg

In the previous post, CCC C960 Blitz Championship (October 2021), I wrote,
Given that engines' evaluations for every move are available in the event's PGN game scores, perhaps there is something to be learned about the 960 different start positions. That investigation would make a good follow-up post.

Make that two good follow-up posts. The first post was on my main blog, Evaluating the Evaluations (November 2021), where I concluded,

Now that I have a tool for rapidly evaluating the engine evaluations, what can I do with it? The first task will be to put it to work on the 960 start positions used in chess960.

The second post is this one. I had already downloaded a few PGN files from recent engine vs. engine events, so the first question was which one to use. I decided to continue with the games from an event that I covered earlier this year in another post on this blog, TCEC C960 FRC3 (March 2021). At that time I noted,

Except for an occasional CCRL game, I can't remember ever looking at an engine vs. engine chess960 game. Is there anything to be learned from such an exercise, or is the play of the engines beyond comprehension?

TCEC FRC3 was a 50 game match won by KomodoDragon over Stockfish on a final score of +2-1=47. The seven mandatory tags in the PGN header for the first game look like this:-

[Event "TCEC Season 20 - FRC3 Final"]
[Site "https://tcec-chess.com"]
[Date "2021.03.14"]
[Round "1.1"]
[White "KomodoDragon 2671.00"]
[Black "Stockfish 20210226"]
[Result "1/2-1/2"]

I loaded the file into my database, added the concept of SP, and produced the following chart. It covers the first 22 games of the match. Each start position (SP) was played twice, where KomodoDragon always had White in odd-numbered games. In a match between humans, this pattern would risk giving an advantage to one of the players, but in games between engines, it's harmless.

The last two columns show the first move, as chosen by White, and the value ('wv') calculated by the engine for that move. I could have also shown the principal variation ('pv') calculated by White, but that wouldn't add much to an initial understanding of the data. The same data is available from the PGN file for all moves by both sides in a game.

Since the data covers only the first move of 25 SPs (50 games) out of the full set of 960 SPs, it's obviously just scratching the surface. Suppose we had data for the first few moves of all 960 SPs from many different engines played over a long period of time. What might we learn from this? I would want an answer to that question before spending too much effort collecting more data.

30 October 2021

CCC C960 Blitz Championship

There's one more idea left from the recent post, Crossover Ideas from my Main Blog (October 2021):-
Review the recent CCC chess960 tournament • The semifinal finishes this weekend. Can we expect a final? Short answer: Probably.

Change that 'Probably' to 'Yes'. In the most recent post in the ongoing TCEC/CCC engine saga, TCEC Cup 9, CCC C960 Blitz Final : Both Underway (October 2021), I continued,

In the 'Chess960 Blitz Semifinals', Stockfish finished a point ahead of Dragon as both engines qualified for the final match. Only one game of their 40-game [semifinal] minimatch was decisive, with Stockfish winning. Lc0 lost three games to each of the two engines, winning none. The other three engines were far behind.

In the final match Stockfish beat Dragon +10-1=589. Yes, more than 98% of the final games were drawn. Earlier this year, in TCEC C960 FRC3 (March 2021), I reported,

In the 'FRC 3' final, KomodoDragon beat Stockfish by a score of +2-1=47. A 94% draw rate echoes the sort of result we expect from a traditional chess match (SP518 RNBQKBNR) between engines.

Note that the CCC's Dragon and the TCEC's KomodoDragon are the same engine. It's also worth noting that Stockfish switched to NNUE evaluation last year, while Dragon is also an NNUE engine, as I noted a year ago on my main blog in Komodo NNUE (November 2020). Is the high percentage of draws because they both use the same technology for evaluating positions?

The following chart shows the result of the CCC semifinal round. Stockfish and Dragon finished 1st and 2nd, ahead of 3rd place Lc0 and three other engines. I know the black background makes the chart hard to read, but the individual game results, especially the losses in red, are clearly discernible.

Stockfish didn't lose a single game during the event, while Dragon lost only one game, to Stockfish. As mentioned above, both engines beat Lc0 three times, which itself lost only a single game to the three engines in the bottom half of the crosstable. The bottom half is a sea of red.

Given that engines' evaluations for every move are available in the event's PGN game scores, perhaps there is something to be learned about the 960 different start positions. That investigation would make a good follow-up post.

16 October 2021

Crossover Ideas from my Main Blog

Since this is a month with five Saturdays, I get three opportunities for a chess960 post. By coincidence, I have exactly three ideas for those posts.

1) Add CFAA posts to the download tag

A few months ago I created New Label 'Download' (August 2021), to keep track of posts with a download, most likely a PGN file. Before starting this blog I used my main blog 'Chess for All Ages' to write about chess960. Were there any posts on that blog to add to the 'Download' label? Short answer: No.

2) Review Carlsen's chess960 activity

Again referring to my main blog, I've been building a reference for World Champion Magnus Carlsen's playing record over the past three years. The most recent post was Carlsen's TMER 2019-21, 'Online = Y' (October 2021), where TMER stands for 'Tournament, Match, and Exhibition Record'.

So far I've identified two chess960 events for the TMER. Were there others? Short answer: Yes. A post on this blog, Carlsen Wins Lichess Again (March 2019), discussed one chess960 event and pointed to a previous event, neither of which is listed on the TMER. While researching those two events, I discovered a third. This needs more work.

3) Review the recent CCC chess960 tournament

Another recent post on my main blog, TCEC Testing Cup 9; CCC C960 Blitz Semifinal (October 2021), refers to Chess.com's ongoing 'Computer Chess Championship' (CCC), where the latest event is the 'Chess960 Blitz Championship'. The semifinal finishes this weekend. Can we expect a final? Short answer: Probably.

I'll come back to Carlsen's chess960 activity and the CCC chess960 tournament in the next two posts scheduled for this month. I expect both of those posts will lead to new ideas. That's life in the chess960 blogosphere!

27 March 2021

TCEC C960 FRC3

On my main blog, where I've been tracking the world's two foremost, ongoing engine vs. engine competitions, the most recent fortnightly post, TCEC FRC3, CCC Rapid 2021 : Both Finals Underway (March 2021), noted,
FRC3 [Fischer Random Chess 3] has reached the final match, where Stockfish and KomodoDragon are tied with one win each after 29 of the 50 games have been played. The following chart from the TCEC Wiki shows the different stages of the event.

Here's the chart mentioned in the quote, taken from TCEC FRC 3 (wiki.chessdom.org):-

I ended the post on my main blog saying,

After the final match finishes, I'll have more to say about the tournament on my chess960 blog. I covered the previous edition in TCEC FRC2. In that event, Stockfish beat LCZero +8-0=42.

At the time I wrote the referenced post TCEC FRC2 (December 2020) -- which incorporated the same TCEC flow chart shown above -- the TCEC treated chess960 as a second class citizen, reported only in a section of the TCEC Wiki's home page. Since then, the main page section has been promoted to a separate page, TCEC FRC 2 (wiki.chessdom.org), which announces,

Though at the time advertised as [a] Bonus event, the TCEC Fischer Random Chess will as of now be a regular part of seasonal events.

According to the Wiki's history page, the move happened exactly a month ago:-

23:07, 27 February 2021 [...] (moved from Main page)

'Will be a regular part of [TCEC] seasonal events' -- that's one small step for TCEC, one significant step for chess960. The TCEC Wiki's 'FRC 3' page includes a chart showing the scores of the event's four stages.

A similar chart is now available on the Wiki's 'FRC 2' page. In the 'FRC 3' final, KomodoDragon beat Stockfish by a score of +2-1=47. A 94% draw rate echoes the sort of result we expect from a traditional chess match (SP518 RNBQKBNR) between engines. Will the TCEC FRC organizers be forced to dictate the chess960 opening variations, just as they do for SP518? Let's hope not.

Except for an occasional CCRL game, I can't remember ever looking at an engine vs. engine chess960 game. Is there anything to be learned from such an exercise, or is the play of the engines beyond comprehension? Watch this space; if not for FRC3, maybe for FRC4 or beyond.

[The title of my 'TCEC FRC2' post on this blog is nearly identical to the title on the TCEC Wiki, 'TCEC FRC 2'. To avoid confusion in future reports on the TCEC C960 events, I decided to change the title on this current FRC3 post.]

26 December 2020

TCEC FRC2

On my main blog I've been keeping track of the TCEC engine vs. engine tournaments. Last month, in Stockfish Wins TCEC Cup 7; CCC GPUs Back (November 2020), I reported,
The [TCEC] '!next' plan says, 'next FRC2 testing and FRC2 ~1.5 weeks'. When was FRC1? As far as I can tell, it was more than six years ago. [...] I'm looking forward to reporting on FRC2 for [my chess960] blog.

Two weeks later, in TCEC FRC2 Underway; CCC 'Currently Uncertain' (November 2020), I reported,

After 'Sufi Bonus 3', the [TCEC] ran a chess960 event, dubbed 'FRC2'. It started with 16 engines in four 'Leagues' (A to D), followed by eight engines in two 'Semileagues' (1 to 2), followed by four engines in a 'Final League', followed by two engines in a 'Final' match. The 'Final League' is currently underway.

Another two weeks passed and in TCEC S20 Underway; CCC Less Uncertain (December 2020), I reported,

In the FRC2 Final League, LCZero and Stockfish finished first and second to qualify for the 50-game final match. Stockfish beat LCZero +8-0=42.

The first of the three posts above linked to Stockfish, the Strong (July 2014) on this blog, plus two other followup posts based on FRC1. FRC1 was held three and a half years before AlphaZero made waves with its revolutionary AI/NN technology, soon to be followed by Leela Chess Zero (aka LCZero / LC0). The chart below overviews the different events that made up FRC2. The top portion of the chart flows upward; the bottom portion flows downward.


Source: TCEC Wiki

The semifinal event, dubbed 'Final League' in TCEC nomenclature, had Komodo representing the traditional engines that competed in FRC1, Lc0 and AllieStein representing the AI/NN generation of engines, and Stockfish representing the even newer NNUE generation. I haven't decided if I'm going to spend time looking at the games from FRC2. We already have years of engine experience documented in the CCRL datasets (see the right sidebar under 'Resources') and I'm not sure what can be gleaned from the latest TCEC experiment.

27 June 2020

No Quitting Here!

In last week's post, Chess960 on Playchess.com, I responded to a quote from IM Sagar Shah with:-
Sagar Shah says, 'That's the thing in chess960. If you're not careful you can quickly run into a lost position.'

Since I'm something of a specialist for running into lost chess960 positions, I should document some of my most painful experiences. But which games should I choose? There are so many of them.

If I had any common sense I would stop this chess960 blog here and now. It's been five years to the day since I signalled my first attempt to quit in Whispering a Fond Adieu! (June 2015; 'Bye for now! - Mark'). I managed to stay away for 18 months, then came roaring back with 'Everyone I Know Plays Chess960' (January 2017). The title of that post was a quote from GM Peter Svidler where the complete thought was, 'Everyone I know plays chess960 with great pleasure.' Copy that!

Yes, chess960 continues to be a great pleasure for me. I started playing on correspondence servers in 2008 -- Chess960? I'm Hooked! (September 2008) -- and am still hooked. According to my records, I played 92 games on Schemingmind, most recently in 2016. I played another 136 games on LSS, where I currently have a half-dozen games underway.

With more than 200 games under my belt, I have plenty of examples to choose from -- wins, losses, and draws -- most of which were analyzed fairly deeply while they were being played. All of the LSS games were played with the help of an engine, so I'll start with those. After 12 years of playing chess960, I still don't understand much about its opening theory, making it a logical area to focus on.

So here's the plan: I'll continue to post twice a month. One post will be to keep up with any news; one post will be to learn something about opening theory. Maybe I'll eventually discover how to avoid running into a lost position.

30 December 2017

Process Improvements

Continuing with Engine Trouble (September 2017; 'Talk about a disastrous tournament!'), in that post I described the background with a few sentences:-
The tournament was the final section of the LSS 2015 Chess960 Championship, a three-stage elimination tournament. [...] The 2016 final tournament starts soon, but I probably won't participate. I've taken enough psychological punishment for one year.

I finally decided to participate in the 2016 final for two reasons:-

  • It's not so easy to reach the final and this might be my last opportunity.
  • I couldn't do any worse than in the 2015 event.

I also decided not to upgrade my engines for this event, but I did introduce a number of 'process' improvements in the way I use engines. Specifically, I tried four new techniques that I'll discuss separately:-

  1. Castling manually
  2. Using two engines with different qualities
  3. Making first a null move, then applying opponent's expected move
  4. Using coarser granularity

1. Castling manually

In another follow-up post to 'Engine Trouble', The Seeds of Disaster, I noted, 'A chess960 opening can thus be logically divided into three phases': before either player has castled, after one player has castled, and after both players have castled. (The phrase 'player has castled' can also mean that a player has somehow lost the castling privilege by moving the King or by moving both Rooks.) In many games I've noticed some apparent inconsistencies in engine evaluation across these three phases. Was this another problem related to tapered evaluation as discussed in Chess Engines : Advanced Evaluation (September 2015)?

I decided to experiment by evaluating two identical lines. In one line I castled normally; in the other I castled artificially by moving the King and the Rook (in chess960, sometimes only one or the other moves) one square at a time, inserting null moves for the opponent, eventually reaching the castled position. I once used a similar technique in The Engines' Value of Castling (May 2015). I indeed recorded some differences, although it's too early to report on the results because the games are still running.

2. Using two engines with different qualities

The previous discussion is only relevant to chess960, but the remaining points are also relevant to traditional chess. I often use two (or more) different engines to evaluate the same position, then analyze any differences between them. This helps me understand unclear positions, especially where there is unbalanced material on the board. For example, I've often noticed that Houdini is better than Komodo at tactical play, but Komodo is better than Houdini at positional play.

I avoided using Stockfish for this sort of comparison because its search depths don't compare to the other engines. Roughly speaking, it takes Houdini and Komodo twice the time to calculate each successive ply, but Stockfish takes only 50% more time. I was reminded of this during the latest TCEC season, which I wrapped up with Houdini, Komodo, Stockfish, and AlphaZero. In a TCEC report from last month, Interview with Robert Houdart, Mark Lefler and GM Larry Kaufman (chessdom.com), the following exchange took place:-

Nelson [TCEC organizer]: What quality of your program do you think may be superior to your opponent in the Superfinal?

Larry [Team Komodo]: Basically, we have much more comparison with Stockfish because Stockfish is open source so we can easily compare our ideas and see what works better or worse. I don’t really know the inside workings of [Houdini], but what I can tell you is that my belief is that Komodo is better in most things than Stockfish. But there is something holding us back that has to do with search depth. We’ve been trying to figure it out for years, I don’t know what it is, but there is some reason we are not able to get the same search depth as Stockfish even if we tried to copy all their algorithms. We’ve tried experiments where we’ve tried to make Komodo act like Stockfish but it doesn’t work, and I don’t know why, but I feel that if we ever figure that out we’ll just be clearly #1. But almost every time we tried any idea from Stockfish in Komodo, nine times out of ten it makes Komodo weaker.

If the developers of a world class chess engine can't explain the phenomenon, what hope is there for the rest of us? It probably has something to do with pruning, as in Chess Engines : Pruning (September 2015). Long story short, I started comparing Houdini's lower-depth evaluations with Stockfish's higher-depth, but am not yet sure what I'm seeing.

3. Making first a null move, then applying opponent's expected move

It sometimes happens that no matter what move the engine proposes, the expected response by the opponent is always the same. When this situation occurs, I started using a technique where I first make a null move, then apply the expected response, then evaluate the resulting position. Any further analysis of the subsequent variations requires inserting a null move for the opponent. This should help to understand the tradeoffs for the immediate move.

4. Using coarser granularity

I also started experimenting with relaxed granularity. Engines normally return their (numerical) evaluations in units of centi-Pawns (0.01). I've often noted that it's shortsighted to favor one move over another simply because the first has a value of 0.02 and the second has a value of 0.01. This is even more shortsighted when one of the values is 0.00, which can mean all sorts of things. Any relaxing of strict numerical order is what engine developers call 'contempt' (see that Chessdom.com interview for more about the concept), but I started applying it to the moves suggested by the engines. I'll try to cover this in another post.