行ってきた。朝一で奈良入りでもいけなくはなかった気がしないでもないが、せっかくだしということで奈良に2泊3日してきた。奈良の観光の話でもいいけど、流石にどこも回れなかったのでワークショップの話。
前日の夜にAlexからスケジュールの変更のメールが届く。発表者二人が飛行機に乗り遅れたので、その二人を最後に回すため。結局間に合わずWillとKathyによりカラオケスライドメソッド(後述)が発動した。
ホテルから徒歩15分という距離なのに9時15分のオープニングに間に合わず。油断しすぎでした。
【招待講演1】
プログラムの正当性を検証するプログラムにLispを使ったという話(だと思う)。スケジュールが移動して自分の発表が一発目になったのでそれに意識が行き過ぎてたから話が頭に入らなかった…これ聞きながら、今年のはこんなにアカデミックなのかぁとものすごく緊張した記憶しかない。
【自分の】
そのうちスライド挙げます。単にライブラリの紹介。
【Nash: a tracing JIT for Extension Language】
GuileにトレーシングJITを実装したという話。プログラムで使用される時間の大部分はタイトなループによって発生しているということに注目したJITらしい。本家にマージされるかは不明らしい。されてほしいような、されてほしくないような(Sagittariusが選ばれる可能性が下がるという意味で)。
【Ghosts in the machine】
現在のプログラミング環境はコンピュータ黎明期と比較して劣っている、ということをClojureのREPLを例にして示す話。時間内に終わらなかった発表一つ目。休憩中に発表者に質問してどこを目指しているのか質問して聞いたりしていた。視点としては面白いし、一意見として終わらせるには惜しい気がするけど、何から始めるといいんだろう?となるくらいには壮大なビジョンだった。
【R7RS update】
AlexによるR7RSの近況。まだ見ぬSRFI-142とSRFI-143が並んでいたのでArthurで止まっているのか、Johnがまだ提出してないのかは謎。Red Docketは終わったらしいのだが、正式な発表はまだされてない気がするけど、c.l.sで見逃したかな?
【招待講演2】
Guixという比較的新しいパッケージマネージャの話。インストール履歴をリビジョンで管理しているのでうっかり何かを壊しても動くリビジョンに戻せるというのが他のパッケージマネージャとの大きな違いかな。後はGuileで書かれているのでパッケージの定義もS式(Guileで動くプログラム)というのも大きな特徴かな。
【Function compose, Type cut, And the Algebra of logic】
発表者のコミュニケーション能力があまり高くなかったので何が何だかさっぱり分からなかった。ただ、発表資料自体は面白いことが書いてあったので、論文を後で読むことにする。
【Multi-purpose web framework design based on websocket over HTTP Gateway】
SagittariusにWebsocketを入れたこともあり一番気になってた発表なんだけど、今一発表が雲を掴む様な感じでよく分からなかった。また、コミュニケーションの問題も多少あり、納得のいく答えも得られず。後で論文を読む。
ここからカラオケスライドメソッドが発動する。ちなみに、カラオケスライドとは自分が作ったスライドではないものを発表するもと思えばよい。命名はWill。日本のカラオケが流れてくる文字を歌うことに引っ掛けた上手い名前だと個人的には思った。
【miniAdapton: A Minimal Implementation of Incremental Computation in Scheme】
AdaptonというMemoriseのテクニックに一つをSchemeで実装した話。入力ツリーの一部が変更されても共通部分は再利用しているような感じだった。リポジトリも公開されてるし、ソース見た方が早いかな。しかし、代理の発表(論文の共著者であるが) なのにえらくうまいことやるなぁと前回見た時も思ったなぁ。あれくらいうまくしゃべれるようになりたい(なんか違う)。
【Deriving Pure, Functional One-Pass Operations for Processing Tail-Aligned Lists】
リストの共通サフィックス部分を調べるスクリプトをエレガントに書いたという話。ベンチマークがないのでこれが実際にどこまで効くのかは微妙。ナイーブな実装、エレガントな実装ともにO(N)なので、コンスタントの部分だけなのだが、正直論文の実装だとヘルパー手続きの作成にかかるコストの方が高い気がしないでもない。まぁ、コードは論文に載っているので手元でベンチマークとってJasonにメール投げてみればいい話な気がする。
どうでもいいのだが、僕が喋る英語はオランダ語訛りらしい(Arthur談)。どうも子音の発音がネイティブに比べて強いのだそうだ。普通母語に引っ張られるが、第三言語(第二が英語)に引っ張られるのは珍しいんじゃないの?という話をしていた。確かに寡聞にして聞かない話ではある。
Syntax highlighter
2016-09-19
2016-08-30
ログ
そういえばSagittariusにはずっとログを吐き出すためのライブラリがないということをふと思い出した。附属させてるライブラリがログを吐き出すのはさすがにどうかと思うのであんまり考えてなかったのだが、Paellaみたいなのが何のログも吐かないというのはいろいろ面倒だなぁと思ってはいた。
あまりいいデザインというのも思い浮かばないんだけど、なんとなくこんな感じのロガーがあればいいかなぁと思い30分くらいで作ってみた。こんな感じで使える。
とりあえずこれを適当に使ってるスクリプト等に組み込んで使用感を確かめていくことにしよう。
あまりいいデザインというのも思い浮かばないんだけど、なんとなくこんな感じのロガーがあればいいかなぁと思い30分くらいで作ってみた。こんな感じで使える。
(import (srfi :18) (util logging) (util file))
(define (print-log logger)
(trace-log logger "trace")
(debug-log logger "debug")
(info-log logger "info")
(warn-log logger "warn")
(error-log logger "error")
(fatal-log logger "fatal")
(terminate-logger! logger))
(print-log (make-logger +info-level+ (make-appender "[~l] ~w4 ~m")))
(print)
(print-log (make-async-logger +debug-level+
(make-appender "[~l] ~w4 ~m")
(make-file-appender "[~l] ~w4 ~m" "log.log")))
(print (file->string "log.log"))
#|
[info] 2016-08-30T16:02:29+0200 info
[warn] 2016-08-30T16:02:29+0200 warn
[error] 2016-08-30T16:02:29+0200 error
[fatal] 2016-08-30T16:02:29+0200 fatal
[debug] 2016-08-30T16:02:29+0200 debug
[info] 2016-08-30T16:02:29+0200 info
[warn] 2016-08-30T16:02:29+0200 warn
[error] 2016-08-30T16:02:29+0200 error
[fatal] 2016-08-30T16:02:29+0200 fatal
[debug] 2016-08-30T16:02:29+0200 debug
[info] 2016-08-30T16:02:29+0200 info
[warn] 2016-08-30T16:02:29+0200 warn
[error] 2016-08-30T16:02:29+0200 error
[fatal] 2016-08-30T16:02:29+0200 fatal
|#
ロガーがログレベルと同期をコントロールして、アペンダーが実際にログを吐き出す。どこかでみたことあるようなモデルではある。まぁ、本職はその言語を使うのでなんとなく似通ったのだろう。make-appenderでつくられるアペンダーは標準出力(正確にはcurrent-output-portが返す値)に吐き出す。アペンダーの第一引数はログのフォーマット。現状手の込んだ出力はできないが、当面はこれでも問題ないだろう。ちなみに、~wの後ろに続く4はSRFI−19で定義されているフォーマットの一つ。現在は一文字しかみないが、そのうちなんとかするかもしれない。make-async-loggerはスレッドを一つ消費する代わりにログの書き出しをバックグラウンドでやってくれる。アペンダーが複数あったり、処理が重い(メール送るとか) 等あるときに役に立つと思う。とりあえずこれを適当に使ってるスクリプト等に組み込んで使用感を確かめていくことにしよう。
2016-08-29
Ephemeron and reference barrier
Recently, there's the post on SRFI-124 ML about reference barrier. The SRFI doesn't require
Suppose this assumption is correct, then
reference-barrier procedure due to the non-trivial work to implement. Now, there's a post which shows how to implement portably, like this:
(define last-reference 0)
(define (reference-barrier x)
(let ((y last-reference))
(set! last-reference x)
y))
It seems ok, but is it? The SRFI says like this:
This procedure ensures that the garbage collector does not break an
ephemeron containing an unreferenced key before a certain point in
a program.
The program can invoke a reference barrier on the key by
calling this procedure, which guarantees that
even if the program does not use the key, it will be considered
strongly reachable until after reference-barrier returns.
Due to the lack of knowledge and imagination, I can't imagine how it should work precisely on like this situation:
(let loop ()
(when (some-condition)
(reference-barrier key1)
(reference-barrier key2)
(do-something-with-ephemerons)
(loop)))
Should key1 and key2 be guaranteed not to be garbage collected or only the last one (key2 in this case)? If the first case should be applied, then the proposed implementation doesn't work since it can only hold one reference. To me, it's rather rational to hold multiple reference since it's not always only one key needs to be preserved. However if you extend to hold multiple values, then how do you release the references without calling explicit release procedure? If the references would not be released, then ephemerons would also not release its keys. Thus calling reference-barrier causes eternal preservation.Suppose this assumption is correct, then
reference-barrier should work like alloca with automated push so that the pushed reference is on the location where garbage collector can see them as live objects. In this sense, allocated references are released automatically once caller of reference-barrier is returned and multiple references can be pushed. Though, it's indeed non-trivial task to implement. I hope there's no errata which says reference-barrier is *not* optional, otherwise it would be either a huge challenge or dropping the support...
2016-08-24
暗黙の総称関数 (部分解決編)
昨日の続き
総称関数が暗黙的に大域定義されるという話なのだが、昨日のアイデアを元に直してみた。現在のHEADでは以下のようにしても大域には定義されない。
問題は
これとは直接は関係ないのだが、変数名の重複をチェックするようにした。以下のようなものが動なる。
完璧な解決ではないが、メソッドの局所定義とか(個人的には)やらないだろうし、9割くらいの混乱を未然に防げるんじゃないかなぁ。
総称関数が暗黙的に大域定義されるという話なのだが、昨日のアイデアを元に直してみた。現在のHEADでは以下のようにしても大域には定義されない。
(import (rnrs) (clos user)) (let () (define-generic foo) foo) foo ;; -> &undefined
define-genericは今まで無条件に大域に束縛を挿入していたが、現在はトップレベルで定義された際のみとしている。これは別に難しいことではなかったので割愛(実現するのに他の不具合を直す必要があったが)。問題は
define-methodである。現状ではほぼうまく動くのだが、以下のような使い方をすると動かない。
(import (rnrs) (clos user)) (let ((foo 'foo)) (define-method foo (o) o) (print (foo 1))) ;; -> error
define-methodが局所定義された際には局所定義された総称関数をまず探し、なければ大域定義されたものを探すという手順を取っている。そして両方とも見つからなかったらdefine-genericを挿入する。ちなみにここで挿入されたものは局所定義になるので、大域には定義されない。上記の例がうまく動かないのは、メソッド名と局所変数名が同一だから。マクロ展開は当然コンパイル時に行われるので、同名の変数を探すことはできるが、その変数が何を指しているかというのは、指しているものがリテラルである場合を除き、実行時まで分からない。この使い方をすることはそうないだろうと踏んで、現状ではこれで妥協することにした。何か妙案が浮かんだ時にまた直すかもしれない。(どうでもいいが、投げられるエラーが&assertionなのはおかしい気がしてきた。少なくとも&errorじゃないと…)これとは直接は関係ないのだが、変数名の重複をチェックするようにした。以下のようなものが動なる。
(let ((a 1) (a 2)) a) ;; -> error
(let (("a" 1)) 1) ;; -> error
もともと単にサボっていただけなのだが、総称関数が以下のように定義された際に混乱を招きそうだったので:
(let () (define foo) (define-generic foo) (define-method foo (o) o) (foo 1)) ;; ???そもそもエラーなので何が起きても問題ないんだけど、変な混乱を招くよりはコンパイラに怒られた方がいろいろ楽な気がしたというのが本音。
完璧な解決ではないが、メソッドの局所定義とか(個人的には)やらないだろうし、9割くらいの混乱を未然に防げるんじゃないかなぁ。
2016-08-23
Implicit generic function creation
define-method would create a generic function implicitly if there's none defined yet. It's convenient and I've written libraries (probably only one) depending on this behaviour. (c.f. (binary data))Convenience would usually be a trade-off of consistency, at least in my case. For example, this is a long standing bug (though, I've just issued):
(import (clos user)) (let () (define-generic foo)) foo ;; -> should be &undefinedThis is because
define-method would create a generic function during macro expansion, and it would be an unexpected result in this case if it didn't:(begin (define-generic foo) (define-method foo (o) #t))In this case,
define-method should not make implicit generic function, but macro expansions are done in the same compilation time as define-generic. Thus, it's impossible to know if it should create or not. To let define-method know it shouldn't, define-generic also inserts binding into current environment (a library) during macro expansion.Can't I do better? Creating global binding during macro expansion is rather ugly, but I don't I can get rid of it (or maybe the future?). Yet, I think I can avoid to create unwanted one like above example. It's still just an idea but if
define-generic and define-method can see if they are used in a scope, then it seems there's a way. Since Sagittarius has current-usage-env and current-macro-env procedures, it is possible to access compile time environment during macro expansion. Thus, it should even be able to detect whether or not define-method should create an implicit generic function or not.This should work, let's see.
2016-08-22
Cトランスレータ
ふとスタンドアロンのバイナリができるといいかなぁと思ってこんなものを書いてみた(使うには0.7.8のHEADが必要)。以下のように使う:
なんでこんなものを書いたかというと、特に理由はないんだけど、バイナリ一個で動くといいかなぁと思ったから。今のところランタイムとして最低でも DLL か .so (OSXなら .dylib) がいるがこの辺はそのうち気が向いたら何とかする予定。
キャッシュをそのままダンプしているので、性能的なメリットは一切ない。強いて言えば多少ファイルアクセスが減るくらい。今のところ
作っておけば気が向いたときに改善されていくのではないかというメソッドなので、しばらく実用にはならない可能性が高い。
$ ./scheme2c -o out.c foo.scm $ gcc `sagittarius-config -L -I -l` -O2 out.c中身はほぼなんでもいいんだけど、引数を受け取るにはSRFI-22の
mainがないとたぶんうまいこと使えないはず(command-line手続きで引数が取れない)。 ちなみに出来上がりCファイルは超巨大になる可能性がある。参考までに、このスクリプト自信をCに変換したら45MBのファイルになった(64bit環境)。ファイルが巨大になるのはこのスクリプトが何をやっているかを見ればすぐにわかるのだが、端的に言うとキャッシュファイルを引っ張ってきてCのバイト列にしてるから。なんでこんなものを書いたかというと、特に理由はないんだけど、バイナリ一個で動くといいかなぁと思ったから。今のところランタイムとして最低でも DLL か .so (OSXなら .dylib) がいるがこの辺はそのうち気が向いたら何とかする予定。
キャッシュをそのままダンプしているので、性能的なメリットは一切ない。強いて言えば多少ファイルアクセスが減るくらい。今のところ
sagittarius-configは Windows 版にはつけてないので VC でバイナリを作りたい場合は多少の工夫が必要。ちなみに、インストール時にパスが決まるからインストーラでやらないといけないというのがついてない理由。作っておけば気が向いたときに改善されていくのではないかというメソッドなので、しばらく実用にはならない可能性が高い。
2016-08-11
不要レジスタの除去とSchemeコード
VMにレジスタを追加するとスレッドセーフなパラメタと同様な感じで使えるということがあり、SagittariusではVMにいくつかマクロ展開器用のレジスタを持たせていた。あまりいい手ではないし、無駄にVMのサイズを増やすことになるのでいつかは何とかせねばと思いつつもかなりの期間放置していた。っが、最近いい方法を思いついたのでえいや!と除去。VMのサイズが5ワード減って752バイト(64ビット環境)になった。焼け石に水もいいところである。
【やったこと】
基本的には非常に簡単で、VMのレジスタをScheme側でパラメータ化してそれらを使っていたCの実装を全てScheme側に移動させただけ。Scheme側といってもプレビルドされているものなので、本体サイズが小さくなるということは残念ながらない。単にVMのサイズが多少減るのと、自己満足度が多少上がっただけである。
言うは易し行うは、キャッシュのせいで、多少難かった。除去したVMレジスタはマクロ展開器ようにしぶしぶ追加したもの。こいつのせいでマクロのキャッシュが無駄に複雑になっていた(今でも不必要に複雑だが)。SchemeのパラメータをCから呼び出すということは可能な限り避けたいと思ったので、マクロ展開器に関するコードをごそっとScheme側に移動させる必要があった。そうすると、Cで作られた展開器と密接につながっていたキャッシュの読み出しと書き出しが壊れる。キャッシュ機構自体が無駄に複雑なので面倒なデバッグに突入。後は気合でって感じだった。
これによるパフォーマンスの低下があるかなぁと思ったけど、顕著に表れるようなものはなかったのでよし。
【Schemeコードの混在】
VMのレジスタにはリーダーマクロやR7RSの
Schemeコードのプリコンパイル自体はすでにあるのだから、ちょっと手を入れればいけるだろうと思って手を入れてみたら動いたという感じ。残念ながらCで書かれた手続きからSchemeの手続きを呼ぶことはできないし、SchemeからCの関数を直接呼び出すこともできない。それでも、混在可能というのは非常に楽である。依存関係があるので純粋にSchemeで書くより多少気は使うものの(ちなみにマクロは使えない)、Cでごにょごにょやるより遥かに楽である。これを使って
流石にこれはパフォーマンスに影響がでて、例えば
この調子でVMの不要もしくはあってほしくないレジスタを削除していきたいところである。
【やったこと】
基本的には非常に簡単で、VMのレジスタをScheme側でパラメータ化してそれらを使っていたCの実装を全てScheme側に移動させただけ。Scheme側といってもプレビルドされているものなので、本体サイズが小さくなるということは残念ながらない。単にVMのサイズが多少減るのと、自己満足度が多少上がっただけである。
言うは易し行うは、キャッシュのせいで、多少難かった。除去したVMレジスタはマクロ展開器ようにしぶしぶ追加したもの。こいつのせいでマクロのキャッシュが無駄に複雑になっていた(今でも不必要に複雑だが)。SchemeのパラメータをCから呼び出すということは可能な限り避けたいと思ったので、マクロ展開器に関するコードをごそっとScheme側に移動させる必要があった。そうすると、Cで作られた展開器と密接につながっていたキャッシュの読み出しと書き出しが壊れる。キャッシュ機構自体が無駄に複雑なので面倒なデバッグに突入。後は気合でって感じだった。
これによるパフォーマンスの低下があるかなぁと思ったけど、顕著に表れるようなものはなかったのでよし。
【Schemeコードの混在】
VMのレジスタにはリーダーマクロやR7RSの
includeのためにあるものもあった(取っ払ってやった、更に2ワード減ったぜ)。こいつらはコンパイラが依存していたり、load手続きが使用していたりするので、単純にプリコンパイルされたSchemeに放り込むことができなかった。気合でパラメータ関連をCで書いて何とかするという手もなくはなかったんだけど、年を取ると楽な方に逃げたくなる。ということで、スタブファイル内にSchemeコードが書けるようにしてみた。Schemeコードのプリコンパイル自体はすでにあるのだから、ちょっと手を入れればいけるだろうと思って手を入れてみたら動いたという感じ。残念ながらCで書かれた手続きからSchemeの手続きを呼ぶことはできないし、SchemeからCの関数を直接呼び出すこともできない。それでも、混在可能というのは非常に楽である。依存関係があるので純粋にSchemeで書くより多少気は使うものの(ちなみにマクロは使えない)、Cでごにょごにょやるより遥かに楽である。これを使って
(sagittarius)ライブラリ内に簡易パラメータを作り、load周りのVMレジスタをそいつで実装。それに伴って関連するCのコードをごそっと消してやった。すっきり。流石にこれはパフォーマンスに影響がでて、例えば
(rfc http)のように大量に他のライブラリ(数えたら131個あった)に依存するものをキャッシュなしで読み込むと大体10~15%程度パフォーマンスが落ちた。Cでは手続き内でやっていた処理がScheme側に移動したことで手続き呼び出しに変わったこと等によるオーバーヘッドだろうなぁとは思っているものの、まぁしょうがないか。キャッシュになってさえいればほぼ変わらないので(キャッシュはCなので当たり前だが)、初回起動のみが遅いと思えばそこまで気にすることでもないかなぁ。そもそもコンパイルとマクロ展開が遅いでそっちを先に何とかしないとという話である。(text sql)のコンパイルとか環境によっては10秒かかるし…2000行以上にわたるマクロを10秒で展開すると思えばそこまで悪くないともいえるのか?っがコンパイル待ってる間にコーヒー飲み終わる勢いなのは流石になぁ…この調子でVMの不要もしくはあってほしくないレジスタを削除していきたいところである。
2016-08-03
WebSocketクライアント
連投の二つ目
なんとなくやる気とか刺激とか取り戻すためにWebsocketのライブラリを書いてみた。WebSocketはRFC6455で定義されているプロトコル(詳細はRFC読むべし)で、実装もそんな大変そうでもないので書いてみた。
書いてる途中でSaitoAtsushiさんの実装の存在を思い出してライブラリの名前を確認したら、
簡単な使い方。
すでにあるライブラリにぶつけにいったということもあり、例外とかかなり頑張って作っている。こんなに例外階層作って、しかも例外ハンドリングをまじめにやったのってたぶん初めてじゃないかなぁ。ライブラリ自体の構成は割とスタンダードで、ユーザーレベルAPI、中間レベルAPI、低レベルAPIという感じになっている。低レベルAPIはサーバー書くときに便利に使え(現状ではサーバーをどうするかあんまり考えていない)、中間レベルは何かしらプログラム的にやるのに、ユーザーレベルはJavaScriptのWebsocketみたいな感じで使える。
作って2日なので、作りこみが足りない部分はあるかもしれないが、簡単なチャットサーバーとクライアントみたいなのは作れたので紹介してみた。
なんとなくやる気とか刺激とか取り戻すためにWebsocketのライブラリを書いてみた。WebSocketはRFC6455で定義されているプロトコル(詳細はRFC読むべし)で、実装もそんな大変そうでもないので書いてみた。
書いてる途中でSaitoAtsushiさんの実装の存在を思い出してライブラリの名前を確認したら、
(rfc websocket)とまるかぶりだった。Pegasus用のformulaまで書いていただいているのにこのままぶつけるのもなぁと思い確認してみたところ
とのことだったので、遠慮なくぶつけさせていただいた。他の名前の候補としては、@tk_riple 実際のところ、すでに入れている人がどれだけいるのかなぁという気もするんですよね。 本家でやるなら私自身も本家を使うようになる (私の実装はもうメンテされず、相対的に価値が暴落する) ので、そういう「普通の名前」は本家の方で取って欲しいという気持ちがあります。— 齊藤敦志 (@SaitoAtsushi) 2 August 2016
(rfc :6455 websocket)とか(rfc websockets)(Cのlibwebsocketsに倣って)とかあったけど、「普通の名前」(RFCの名前)を使うことにした。簡単な使い方。
(import (rnrs) (rfc websocket)) ;; Creates WebSocket object (define websocket (make-websocket "wss://echo.websocket.org")) ;; Sets event handlers (websocket-on-open websocket (lambda (ws) (display 'CONNECTED) (newline))) (websocket-on-text-message websocket (lambda (ws text) (display text) (newline))) ;; Connects to the server (websocket-open websocket) ;; Sends a message (websocket-send websocket "Hello") ;; Close it (websocket-close websocket)ユーザーレベルAPIは基本的にWebSocketオブジェクトを返すので、以下のようにも書ける。
(websocket-close
(websocket-send
(websocket-open
(websocket-on-text-message
(websocket-on-open (make-websocket "wss://echo.websocket.org")
(lambda (ws) (display 'CONNECTED) (newline)))
(lambda (ws text) (display text) (newline))))
"Hello"))
どっちがいいかは好みだろうけど(流石に下のはあまり使わないか?)。ドラフトのまま絶賛放置中(期限切れてるから破棄されてるの?)のWebSocket over HTTP/2にも頑張れば対応できるようにはしてある(プラグイン書くだけ)。すでにあるライブラリにぶつけにいったということもあり、例外とかかなり頑張って作っている。こんなに例外階層作って、しかも例外ハンドリングをまじめにやったのってたぶん初めてじゃないかなぁ。ライブラリ自体の構成は割とスタンダードで、ユーザーレベルAPI、中間レベルAPI、低レベルAPIという感じになっている。低レベルAPIはサーバー書くときに便利に使え(現状ではサーバーをどうするかあんまり考えていない)、中間レベルは何かしらプログラム的にやるのに、ユーザーレベルはJavaScriptのWebsocketみたいな感じで使える。
作って2日なので、作りこみが足りない部分はあるかもしれないが、簡単なチャットサーバーとクライアントみたいなのは作れたので紹介してみた。
例外ハンドリング
連投の一つ目。
一つ前の記事で
ぼ~っと考えて実装したら動いちゃった系の解決方法で、特に苦労とかなかったんだけど、実装前に気になっていた点が以下:
二つ目はC側のコールスタック。継続の境界は
具体的にはこんな感じで例外が投げられたとする。
このバグのおかげで別のバグをつぶすこともできたし、YpsilonとMoshのバグを発見することもできたのでいいバグ(?)だった。
一つ前の記事で
guardの例外の再送出をwith-exception-handlerで受けると無限ループに陥る問題を解決した話。結論を先に書くと、raise、raise-continuable及びwith-exception-handlerをSchemeで実装して、継続の境界を作らないようにした。ぼ~っと考えて実装したら動いちゃった系の解決方法で、特に苦労とかなかったんだけど、実装前に気になっていた点が以下:
- Cで
with-exception-handlerを使っているか - 使ってなかった。なんでこれC側にあるんだろう状態だった。
- C側で例外投げたらどうなるの
- この記事の肝、気になったら続きを読んで。
(core)ライブラリだけでは使えなくなったけど、これに依存するコードはないと思うのでまぁ、問題ないだろう。二つ目はC側のコールスタック。継続の境界は
Sg_Apply系の関数で作られるんだけど、コールスタックの深いところで例外が投げられたらどうする?ってのをぼ~っと考えていた。はい、longjmpを使うだけでした。具体的にはこんな感じで例外が投げられたとする。
Call stack
+----------+ +---------+ +---------+ +---------+
--| Sg_Apply |--| C func1 |--| C func2 |--| C func3 |
+----------+ +---------+ +---------+ +---------+
^^^^^^^
C error
こんな感じでCのraiseは呼ばれた時点でVMのスタックにSchemeのraiseを呼び出す継続フレームを入れる。
Before
PC = CALL
Stack
+--------------------------+
| Return to previous frame |
+--------------------------+
:
After
PC = RET
Stack
+--------------------------+
| Return to calling raise | --> PC to return = CALL raise
+--------------------------+
| Return To previous frame |
+--------------------------+
:
with-exception-handlerとraise-continuableの組み合わせだと、呼び出し元に戻る必要があるので一つ前のフレームを飛ばすわけにはいかない。この状態にしておいて、longjmpを呼び出し、VMのループを再起動する。もともとVM自体はVMループへのjmpbufを持っているのでそれを使うだけ。割とお手軽に解決できてしまった。このバグのおかげで別のバグをつぶすこともできたし、YpsilonとMoshのバグを発見することもできたのでいいバグ(?)だった。
2016-07-29
Exception handling in C world
The following piece of code goes into infinite loop on Sagittarius.
The problem is related to continuation boundary and exception handlers. If you print the k, then only the very first time you'd see error object with
This happens when you capture a continuation and invoke it in the different C level apply. Each time C level apply is called, then VM put a boundary mark on its stack to synchronise the C level stack and VM stack. This is needed because there's no way to restore C level stack when the function is returned. When a continuation is captured on the C apply which is already returned, then invocation of this continuation would cause unexpected result. To avoid this, the VM checks if the continuation is captured on the same C level apply.
Still, why this combination causes this? The answer is
(import (rnrs))
(with-exception-handler
(lambda (k) #t)
(lambda ()
(guard (e (#f #f))
(error 'test "msg"))))
I couldn't figure it out why this never returned at first glance. And it turned out to be very interesting problem.The problem is related to continuation boundary and exception handlers. If you print the k, then only the very first time you'd see error object with
"msg" after that it'd be "attempt to return from C continuation boundary.". But why?This happens when you capture a continuation and invoke it in the different C level apply. Each time C level apply is called, then VM put a boundary mark on its stack to synchronise the C level stack and VM stack. This is needed because there's no way to restore C level stack when the function is returned. When a continuation is captured on the C apply which is already returned, then invocation of this continuation would cause unexpected result. To avoid this, the VM checks if the continuation is captured on the same C level apply.
Still, why this combination causes this? The answer is
raise is implemented in C world and exception handlers are invoked by C level apply. The whole step of the infinite loop is like this:guardwithout else clause captures a continuation.- When
erroris called, thenguardre-raise and invoke the captured continuation. (required by R6RS/R7RS) - Exception handler is invoked by C level apply.
- VM check the continuation boundary and raises an error on the same dynamic environment as #3
- Goes to #3. Hurray!
- Create a specific condition and when
with-exception-handlerreceived it, then it wouldn't restore the exception handler. [AdHoc] - Let
raiseuse the Scheme level apply. (Big task)
2016-07-12
リモートREPLとライブラリ依存関係
最近ちょっとしたツールを作るのにPaella(に付属しているPlato)を使っているのだが、リモートREPLの問題とリロードの問題が面倒だなぁと思ってきたのでちょっとメモ。
依存関係の子に当たる部分を親が解決できればなんとかなりそうなんだけど、現状の仕組みでは子が親を探すことができても親は子を知る術がない。突き詰めていくと
あるとよさそうなものとして、
リモートREPLの問題
これは非常に簡単で#!read-macro=...のようなのが送れない。問題も分かっていて、#!から始まるリーダーマクロは値を返さないで次の値を読みにいく。マクロはリーダー内で解決される。スクリプトや通常のREPLなら特に問題ないんだけど、それ自体を送りたい場合にはあまり嬉しくない。解決方法はぱっと思いつくだけで以下:- リモートREPL用のリーダーを作る
- 面倒
- リモートREPL用ポートを作って、リードマクロを読ませる
- アドホックだけど、悪くない気がする
ライブラリ依存関係
Platoを使う最大の理由はREPLでリロードができるからなんだけど、ハンドラ(と呼んでいるライブラリ)が依存しているライブラリを変更した際に上手いことその変更を反映する方法がない。ハンドラであればリロードが可能なのだが、依存ライブラリだとロードパスの問題とかも出てきて嬉しくない。手動でloadを呼ぶとか、コード自体をリモートREPLに貼り付けるとか(上記の問題が出ることもあるが)方法がないこともないけど今一面倒。依存関係の子に当たる部分を親が解決できればなんとかなりそうなんだけど、現状の仕組みでは子が親を探すことができても親は子を知る術がない。突き詰めていくと
includeされたファイルの変更は追跡できないとかあるし。あるとよさそうなものとして、
- ライブラリファイルの変更を検知したらリロードする機構
- 依存関係の親が子を知る手段
2016-07-08
Syntax parameters (aka SRFI-139)
Marc Nieper-Wißkirchen submitted an interesting SRFI. The SRFI was based on the paper 'Keeping it Clean with Syntax Parameters'. The paper mentioned some corner case of breaking hygiene with
In 'Implementation' section, the SRFI mentions that this is implemented on 'Rapid Scheme', Guile and Racket. Unfortunately there's no portable implementation. In my very prejudiced perspective, Guile isn't so fascinated by macro, so the syntax parameters is probably implemented on top of existing macro expander (I haven't check since Guile is released under GPL, and if I see the code it may violates the license). If so, it might be able to be implemented on
Without any deep thinking, I've written the very sloppy implementation like this:
This implementation violates some of 'MUST' specified in the SRFI.
datum->syntax, which I faced before (if I knew this paper at that moment!).In 'Implementation' section, the SRFI mentions that this is implemented on 'Rapid Scheme', Guile and Racket. Unfortunately there's no portable implementation. In my very prejudiced perspective, Guile isn't so fascinated by macro, so the syntax parameters is probably implemented on top of existing macro expander (I haven't check since Guile is released under GPL, and if I see the code it may violates the license). If so, it might be able to be implemented on
syntax-case.Without any deep thinking, I've written the very sloppy implementation like this:
#!r6rs
(library (srfi :139 syntax-parameters)
(export define-syntax-parameter
syntax-parameterize)
(import (rnrs))
(define-syntax define-syntax-parameter
(syntax-rules ()
((_ keyword transformer)
(define-syntax keyword transformer))))
(define-syntax syntax-parameterize
(lambda (x)
(define (rewrite k body keys)
(syntax-case body ()
(() '())
((a . d)
#`(#,(rewrite k #'a keys). #,(rewrite k #'d keys)))
(#(e ...)
#`#(#,@(rewrite k #'(e ...) keys)))
(e
(and (identifier? #'e)
(exists (lambda (o) (free-identifier=? #'e o)) keys))
(datum->syntax k (syntax->datum #'e)))
(e #'e)))
(syntax-case x ()
((k ((keyword spec) ...) body1 body* ...)
(with-syntax (((n* ...)
(map (lambda (n) (datum->syntax #'k (syntax->datum n)))
#'(keyword ...)))
((nb1 nb* ...)
(rewrite #'k #'(body1 body* ...) #'(keyword ...))))
#'(letrec-syntax ((n* spec) ...) nb1 nb* ...))))))
)
And this can be used like this (taken from example of the SRFI):
#!r6rs
(import (rnrs) (srfi :139 syntax-parameters))
(define-syntax-parameter abort
(syntax-rules ()
((_ . _)
(syntax-error "abort used outside of a loop"))))
(define-syntax forever
(syntax-rules ()
((forever body1 body2 ...)
(call-with-current-continuation
(lambda (escape)
(syntax-parameterize
((abort
(syntax-rules ()
((abort value (... ...))
(escape value (... ...))))))
(let loop ()
body1 body2 ... (loop))))))))
(define i 0)
(forever
(display i)
(newline)
(set! i (+ 1 i))
(when (= i 10)
(abort)))
(define-syntax-parameter return
(syntax-rules ()
((_ . _)
(syntax-error "return used outside of a lambda^"))))
(define-syntax lambda^
(syntax-rules ()
((lambda^ formals body1 body2 ...)
(lambda formals
(call-with-current-continuation
(lambda (escape)
(syntax-parameterize
((return
(syntax-rules ()
((return value (... ...))
(escape value (... ...))))))
body1 body2 ...)))))))
(define product
(lambda^ (list)
(fold-left (lambda (n o)
(if (zero? n)
(return 0)
(* n o)))
1 list)))
(display (product '(1 2 3 4 5))) (newline)
I've tested on Chez, Larceny, Mosh and Sagittarius.This implementation violates some of 'MUST' specified in the SRFI.
- keyword bound on
syntax-parameterizedoesn't have to be syntax parameter. (on the SRFI it MUST be) - keyword on
syntax-parameterizedoesn't have to have binding.
define-syntax-parameterdoes nothingsyntax-parameterizetraverses the given expression.
2016-07-07
Weirdness of self evaluating vector
Scheme's vector has a history of being non-self evaluating datum and self evaluating datum. The first one is on R6RS, and the latter one is R7RS (not sure about R5RS). Most of the time, you don't really care about the difference other than it requires
Have a look at this case:
Now, how about this case?
Back to the first case. The first case sometimes bites me when I write/use R6RS macro in R7RS context. For example, SRFI-64 is implemented in R6RS macro and using it like this:
I have sort of solution (not sure if I do it or not): Internally, symbol and the identifiers converted from symbols without any context (c.f. not using
I haven't decided how it should be. So for now, just a memo and let it sleep.
' (quote) or not. However, you may need to think about the difference and maybe also think self evaluating causes more trouble. One of the particular case (and this is the only case I think self evaluating vector is evil) is when vector is used in macro.Have a look at this case:
(import (rnrs))
(define-syntax foo
(syntax-case ()
((_ e) e)))
(foo #(a b c))
What do you think how it behaves? The answer is depending on the standard. On R6RS, vectors are not self evaluating data so this should be an error. So you can't complain if you'd get a daemon from your nose. On R7RS (of course you should change the importing library name to (scheme base)), on the other hand, vectors are self evaluating data so this should return the input vector.Now, how about this case?
(import (scheme base) (scheme write))
(define-syntax foo
(syntax-rules ()
((_ "go" (v ...) ()) #(v ...))
((_ "go" (v ...) (e e* ...)) (foo "go" (v ... t) (e* ...)))
((_ e ...) (foo "go" () (e ...)))))
(foo a b c d e)
What would be the expansion result of macro foo? I think this is totally up to implementations (if it's not please let me know). For example, Chibi returns vector of something like {Sc #22 #<Environment 4365836288> () t} (syntactic closure, I think), Sagittarius returns vector of identifier, and Larceny returns vector of symbol t. If you put ' (quote) to the result template then the expansion result should be the same as Larceny returns (though, Chibi still returned a vector of syntactic closure, so this might not be defined, either).Back to the first case. The first case sometimes bites me when I write/use R6RS macro in R7RS context. For example, SRFI-64 is implemented in R6RS macro and using it like this:
(import (scheme base) (srfi 64)) (test-begin "foo") (test-equal "boom!" #(a b) (vector 'a 'b)) ;; FAIL!! (test-end)On Sagittarius, R6RS macro transformer first converts all symbols into identifiers, then syntax information will be stripped only if expressions have quote. Now, SRFI-64 is implemented on the R6RS macro transformer and the vector doesn't have quote. Thus, symbols inside of the vector are converted to identifiers. If it's R6RS, then it's an error. But if it's R7RS, it should be a valid script.
I have sort of solution (not sure if I do it or not): Internally, symbol and the identifiers converted from symbols without any context (c.f. not using
datum->syntax) are theoretically the same. So if compiler sees such an identifier, then it should be able to unwrap it safely.I haven't decided how it should be. So for now, just a memo and let it sleep.
2016-06-23
価値観の違い
立場が違えば価値観は当然違う。良い悪いという話ではなく、そういうものだと思っている。一時間半に及ぶ最終面接(といって良いのかあれ?)でふと思ったこととをつらつら書いてみる。
現在絶賛転職活動中の僕は今週に1回(今日終わった)、来週に2回とえらい密度で面接が組まれている。別にここまで積極的にするつもりはなかったんだけどタイミング的にこうなった、正直辛い。っで、今日あった最終面接はその会社のオーナーとだったんだけど、なかなか面白い意見だなぁと思った。曰く
もう一つ別に引っかかったのは、「半年もすれば会社で一番の開発者になれる可能性がある」というもの。職場で勉強しようとかそういう気持ちはないんだけど、お山の大将になるつもりもなくて、「全力で追いかけても追いつけない」くらいの人がいる職場の方がいいんだけどなぁ、とか。たった一文の中で矛盾する二つの意見があるというのは置いておいて。同じ条件なら後者を選ぶという程度ではあるが。最先端を常に追いかけてるテック企業といってる割には、僕程度が半年で頂点に立てるレベルというのも今一矛盾しているような。
ここからは個人的な被雇用者としてあり方なんだけど:
会社自体は創業20何年だけどオーナーが変わったのが去年らしく、どうにもスタートアップ的なのか体育会系的なのか分からないが妙にこう会社に尽くす人が欲しい的な雰囲気を前面に押し出す感があった(体育会系じゃないスタートアップに失礼か?)。来週までに辞意を今の会社に告げないと開始時期が9月になるという時期でもあるので、オファー出したら即日もしくは週末を挟んでの月曜日に返事が欲しいとか(流石にもう少し待ってもらうことになったけど)すごい勢いで急かされた感もある。
来週の面接の出来次第かなぁ。条件自体は今より多少上がるし。
絶賛転職活動中なので興味があれば声をかけていただけると嬉しいです。オランダでの勤務もしくは完全リモートが絶対条件だけど。
現在絶賛転職活動中の僕は今週に1回(今日終わった)、来週に2回とえらい密度で面接が組まれている。別にここまで積極的にするつもりはなかったんだけどタイミング的にこうなった、正直辛い。っで、今日あった最終面接はその会社のオーナーとだったんだけど、なかなか面白い意見だなぁと思った。曰く
- その会社で絶対働きたいという熱意がいる
- 会社名のタトゥーをいれるくらいとか
- 40時間越えて働いても残業代請求しないとか
- 金が欲しい開発者ならGoogleにでも行け
- 熱意があるやつが欲しいそうな
- 給料が高すぎるとよくないから(全社員共通の)上限がある
- ちなみに上限はAmazonのオファーより1万5千ユーロほど低かった
- そして上限を上げるつもりはないらしい(といわれた)
- オファーを比較して自分の市場価値を探るのは好きではない
- 自分もやってたけど意味無いってさ
- 比較せずさっさと受けろという意味だとは思うけど
- 副業をしなければいけないほど給料が安い
- そうでないのなら、Google並みの給料もらってもやってると思う
もう一つ別に引っかかったのは、「半年もすれば会社で一番の開発者になれる可能性がある」というもの。職場で勉強しようとかそういう気持ちはないんだけど、お山の大将になるつもりもなくて、「全力で追いかけても追いつけない」くらいの人がいる職場の方がいいんだけどなぁ、とか。たった一文の中で矛盾する二つの意見があるというのは置いておいて。同じ条件なら後者を選ぶという程度ではあるが。最先端を常に追いかけてるテック企業といってる割には、僕程度が半年で頂点に立てるレベルというのも今一矛盾しているような。
ここからは個人的な被雇用者としてあり方なんだけど:
- 労働を提供する対価を雇用者に求めている
- 仕事のやりがいと称して対価を下げる行為を嫌っている
- 対価に勝る評価方法はない
- 僕の労働をより評価してくれる雇用者に容易に移る
会社自体は創業20何年だけどオーナーが変わったのが去年らしく、どうにもスタートアップ的なのか体育会系的なのか分からないが妙にこう会社に尽くす人が欲しい的な雰囲気を前面に押し出す感があった(体育会系じゃないスタートアップに失礼か?)。来週までに辞意を今の会社に告げないと開始時期が9月になるという時期でもあるので、オファー出したら即日もしくは週末を挟んでの月曜日に返事が欲しいとか(流石にもう少し待ってもらうことになったけど)すごい勢いで急かされた感もある。
来週の面接の出来次第かなぁ。条件自体は今より多少上がるし。
絶賛転職活動中なので興味があれば声をかけていただけると嬉しいです。オランダでの勤務もしくは完全リモートが絶対条件だけど。
2016-06-18
R7RSレコード
WG2の議論でレコードの健全性について出てる(c.f. record types)。個人的にレコードが暗黙的にアクセサを作るかどうかというのはとりあえずどうでもよくて(R6RSでは明示されなければ暗黙的に作るってなってるし)、健全性のテストコードが問題だった。
R7RS-smallのレコードは基本SRFI-9なのだが、SRFI-9ではコンストラクタタグとフィールドが同一の識別子でないとエラーとなっている。SagittariusのSRFI-9の実装はこの条件を
テストコードではtmpという識別子がコンストラクタタグとフィールドに並ぶ形になる。これらは
R6RS版の
どうしたか?普通にk番目でアクセスするように変更した。昔レコードをCLOSと統合した際に既にあった機能な気がしないでもないんだけど、なんでスロット名で引くようにしたんだろう?当時は気付いていなかったか、実はこの機能は後から外に出されたかのどっちかだろう。
一見するとマクロのバグのようで実は違ったという話。
R7RS-smallのレコードは基本SRFI-9なのだが、SRFI-9ではコンストラクタタグとフィールドが同一の識別子でないとエラーとなっている。SagittariusのSRFI-9の実装はこの条件を
bound-identifier=?でチェックしているので、テストコードもまぁ問題なく動くだろうなぁと思っていた。のだが、そうもいかなかった。テストコードではtmpという識別子がコンストラクタタグとフィールドに並ぶ形になる。これらは
free-identifier=?を満たすが、bound-identifier=?は満たさない(識別子の挿入されるタイミングが違うため)。そのためSRFI-9の実装的には問題ない。何が問題か?R6RS版のdefine-record-typeとrecord-accessorの実装が問題だった。R6RS版の
define-record-typeは構文情報を引き剥がして下請けのレコード手続きに定義されたレコードの情報を渡すため、テストコードが生成するレコードが持つフィールド名は全てtmpになる(これはR6RS的には許容されている)。っで、record-accessorは与えられたレコードのk番目のフィールドの値を返す手続きを返すんだけど、このk番目にアクセスするためにスロット名でアクセスしていたのがまずかった。構文情報の引き剥がされてるので、単なるシンボルの比較になるんだけど、全部同じ名前なので常に最初にヒットしたスロットの値を返す。どうしたか?普通にk番目でアクセスするように変更した。昔レコードをCLOSと統合した際に既にあった機能な気がしないでもないんだけど、なんでスロット名で引くようにしたんだろう?当時は気付いていなかったか、実はこの機能は後から外に出されたかのどっちかだろう。
一見するとマクロのバグのようで実は違ったという話。
2016-06-12
Because it's fun
Couple of days ago, I've had a job interview, and one of the interviewer said something very interesting. I don't remember exact sentence but something like this: If I make a framework for hobby, it's okay. I didn't understand what the purpose of this comment, so I said "if there's no framework then no choice, right?" Then, he said "if you need to use this framework for work, then you need to consider a lot of things such as buffer overflow.". Well, sort of agree and sort of disagree.
The reason why I needed to make loads of framework is basically because nobody would make other than me. If it's major language such as Java, then you just need to google it and find something. But I'm using Scheme, more specifically Sagittarius. Sagittarius, unfortunately, doesn't have many libraries. Of course, I'm trying to write something useful, but there's only one resource so it doesn't increase the number drastically. And even more unfortunate thing is that it's not so popular so there's not many users. If number of users is small, then not many libraries are written. (It's rather my fault because I still think it only needs to fit my hand and didn't advertise much...) Then if you need something, you gotta write it.
The part I agree with the opinion is the cost of coding. If there are well maintained libraries for you purpose, using them would reduce some time to redevelop the same functionality. Especially if the library is mature enough, then it takes a lot of resource to make your own implementation such level. (e.g. Spring Framework or so)
But should it only be like this? If you are a programmer, you want to write something from scratch because it's fun, don't you? I hear almost every year new trend framework. Not sure the actual motivations are, might be unmatched framework, might be just for fun, etc. And most of the time, it has some bugs. What I want to say is not something like new frameworks are buggy, but they can be famous even the initial version is buggy. So it's pity if you don't write/show something useful because it might not be perfect.
The very first version may only confirm your own requirement. I usually write such libraries and just put it on GitHub (you'll find loads of junks on my GitHub repository). After using it, I usually notice what's missing or not considered. If you have choice to write whatever you want to, then don't hesitate to write especially under the reason of "buffer overflow". (It's more for me, though)
The reason why I needed to make loads of framework is basically because nobody would make other than me. If it's major language such as Java, then you just need to google it and find something. But I'm using Scheme, more specifically Sagittarius. Sagittarius, unfortunately, doesn't have many libraries. Of course, I'm trying to write something useful, but there's only one resource so it doesn't increase the number drastically. And even more unfortunate thing is that it's not so popular so there's not many users. If number of users is small, then not many libraries are written. (It's rather my fault because I still think it only needs to fit my hand and didn't advertise much...) Then if you need something, you gotta write it.
The part I agree with the opinion is the cost of coding. If there are well maintained libraries for you purpose, using them would reduce some time to redevelop the same functionality. Especially if the library is mature enough, then it takes a lot of resource to make your own implementation such level. (e.g. Spring Framework or so)
But should it only be like this? If you are a programmer, you want to write something from scratch because it's fun, don't you? I hear almost every year new trend framework. Not sure the actual motivations are, might be unmatched framework, might be just for fun, etc. And most of the time, it has some bugs. What I want to say is not something like new frameworks are buggy, but they can be famous even the initial version is buggy. So it's pity if you don't write/show something useful because it might not be perfect.
The very first version may only confirm your own requirement. I usually write such libraries and just put it on GitHub (you'll find loads of junks on my GitHub repository). After using it, I usually notice what's missing or not considered. If you have choice to write whatever you want to, then don't hesitate to write especially under the reason of "buffer overflow". (It's more for me, though)
2016-06-04
生産性
ツイッターでも呟いたのだが、体力系と瞬発力系の二つのコーディングテストを最近受けた。体力系は現在進行形だが、コーディング自体は要求された機能を最低限満たしたのでよしとしてしまった。(ネタは面白いんだけど、プライベートでJavaを長時間書く気になれんのよ…)
先週の週末から時間が取れるときにやってたんだけど、途中でJavaに嫌気が差してSchemeで書いてその後Javaに戻ったりしたので、この2言語間(正確にはSagittariusとJavaだが)の生産性の違いが書けそうだなぁと思い、愚痴をこめつつ書くことにした。
先に結論を書くと、Javaは極めて冗長に書く必要があるのでプロトタイプ的なものを作るには向かない。(ので、体力勝負系のコーディングテストに持ってこられると非常に面倒。) 成果はGithubかBitbucketに置けと言われたので、晒しても問題ないと判断して晒す(まさかプライベートリポジトリに置けって意味ではないと思うし)。とりあえず以下はなんとなくな比較
Schemeの実装時間にはテーブル構成の考案にHTML(+Javascript)の時間も含むので多少多めではある(Scheme書いてた時間は3時間くらいかなぁ)。Javaの16時間のうち少なくない時間を割いたのがMavanリポジトリを探す作業。2年近く触ってなかったからいろいろ忘れていた。後はSpring、Hibernateの初期設定とか。一回書くと忘れる系のものは都度Google先生にお伺いを立ててたので。
そうは言っても割と大きめな時間の差があると個人的には思う。言語の習熟度とか、ライブラリ習熟度の差なのかもしれないけど、一応これでも職業Java屋歴10年以上あるのでそこまでの差はないとしたいところ(むしろScheme歴は6-7年なのでJava歴より短い)。個人的に最も生産性(ここでは時間のことを指す)に響いたのはREPLの存在だったと思う。SchemeではリモートREPLを使ってサーバに変更を即時反映させて挙動を確認していたのに対し、Javaでは毎回ビルドするという切ない状況だったのは大きい。対話的に確認できることの偉大さというのを改めて実感した気がする。次いで大きかったのはデザインパターンの有無。JavaだとなんとなくGeneric DAOパターンを使わないといけないかなぁという脅迫観念に襲われたので、とりあえずそれを使ったんだけど、このパターンは恐ろしく生産性が低い。あらかじめあるものを使うのならばいいのだが、エンティティ毎にDAO書いてサービス書いてとかやってるもんだから、非常に面倒だった。(この辺IDEがワンクリックでやってくれると違うのかもしれない。)細かくボディーブローのように効いたのはJavaの1クラス1ファイルの仕様や、キーボードから指が離れる瞬間が多いこと。別にEmacs最高とかいうつもりはないけど、こういうところで地味に時間を取られた気がする。
こういう体力勝負のコーディングテストは最低でも自分が好きな言語でやらせてくれないと途中で息切れするなぁと思った。問題は、僕が好きな言語でやると9割方読めないということだろうか。Schemeいい言語だと思うのだが、なんでこうも人気がないのだろう?
先週の週末から時間が取れるときにやってたんだけど、途中でJavaに嫌気が差してSchemeで書いてその後Javaに戻ったりしたので、この2言語間(正確にはSagittariusとJavaだが)の生産性の違いが書けそうだなぁと思い、愚痴をこめつつ書くことにした。
先に結論を書くと、Javaは極めて冗長に書く必要があるのでプロトタイプ的なものを作るには向かない。(ので、体力勝負系のコーディングテストに持ってこられると非常に面倒。) 成果はGithubかBitbucketに置けと言われたので、晒しても問題ないと判断して晒す(まさかプライベートリポジトリに置けって意味ではないと思うし)。とりあえず以下はなんとなくな比較
| Scheme | Java | |
|---|---|---|
| 実装時間 | 6時間 | 16時間 |
| ファイル数 | 6(自動生成含む) | 61(テストファイル含む) |
| 開発環境 | Emacs | Eclipse |
そうは言っても割と大きめな時間の差があると個人的には思う。言語の習熟度とか、ライブラリ習熟度の差なのかもしれないけど、一応これでも職業Java屋歴10年以上あるのでそこまでの差はないとしたいところ(むしろScheme歴は6-7年なのでJava歴より短い)。個人的に最も生産性(ここでは時間のことを指す)に響いたのはREPLの存在だったと思う。SchemeではリモートREPLを使ってサーバに変更を即時反映させて挙動を確認していたのに対し、Javaでは毎回ビルドするという切ない状況だったのは大きい。対話的に確認できることの偉大さというのを改めて実感した気がする。次いで大きかったのはデザインパターンの有無。JavaだとなんとなくGeneric DAOパターンを使わないといけないかなぁという脅迫観念に襲われたので、とりあえずそれを使ったんだけど、このパターンは恐ろしく生産性が低い。あらかじめあるものを使うのならばいいのだが、エンティティ毎にDAO書いてサービス書いてとかやってるもんだから、非常に面倒だった。(この辺IDEがワンクリックでやってくれると違うのかもしれない。)細かくボディーブローのように効いたのはJavaの1クラス1ファイルの仕様や、キーボードから指が離れる瞬間が多いこと。別にEmacs最高とかいうつもりはないけど、こういうところで地味に時間を取られた気がする。
こういう体力勝負のコーディングテストは最低でも自分が好きな言語でやらせてくれないと途中で息切れするなぁと思った。問題は、僕が好きな言語でやると9割方読めないということだろうか。Schemeいい言語だと思うのだが、なんでこうも人気がないのだろう?
2016-05-25
File monitoring on OS X using FSEvents
(sagittarius filewatch) on OS X is using kqueue (2) currently. Using kqueue isn't bad just having couple of limitation such as no capability of directory monitoring (due to my laziness). It is OK on BSD environment since this is the only choice to do it. However, on OS X, there are FSEvents APIs which allow users to monitor filesystem. When I research it, it doesn't require loads of file descriptors nor limit file/directory only monitoring. So I thought this might be a good underlying implementation for OS X and implemented like this.If it ends without any problem, as usual, I don't write any blog post. Yes, there's a huge problem. It doesn't allow me to write
tail emulator. I first thought my implementation has an issue. So I've written this small piece of code to check if it works as my expect.#include <CoreServices/CoreServices.h>
#include <string.h>
#include <stdio.h>
#include <stdlib.h>
static void callback(ConstFSEventStreamRef stream,
void *callbackInfo,
size_t numEvents,
void *evPaths,
const FSEventStreamEventFlags evFlags[],
const FSEventStreamEventId evIds[])
{
FILE *fp = (FILE *)callbackInfo;
char buf[1024];
const char **paths = (const char **)evPaths;
for (int i = 0; i < numEvents; i++) {
while (1) {
int n = fread(buf, 1, sizeof(buf), fp);
fwrite(buf, 1, n, stdout);
if (feof(fp)) break;
}
fflush(stdout);
}
}
int main(int argc, char **args)
{
if (argc != 2) {
fputs("fstail file", stderr);
exit(-1);
}
FILE *fp = fopen(args[1], "r");
fseek(fp, 0, SEEK_END);
CFStringRef s = CFStringCreateWithCString(kCFAllocatorDefault, args[1],
kCFStringEncodingUTF8);
CFArrayRef ar = CFArrayCreate(NULL, (const void **)&s, 1, NULL);
FSEventStreamContext ctx = {0, fp, NULL, NULL, NULL};
int flags = kFSEventStreamCreateFlagFileEvents;
FSEventStreamRef stream = FSEventStreamCreate(NULL, &callback, &ctx, ar,
kFSEventStreamEventIdSinceNow,
0, flags);
FSEventStreamScheduleWithRunLoop(stream, CFRunLoopGetCurrent(),
kCFRunLoopDefaultMode);
FSEventStreamStart(stream);
CFRunLoopRun();
return 0;
}
This didn't work like tail command, unfortunately. It didn't receive any event after it got first event. If you get a technical problem, you probably search Google or Stack Overflow. Yes, I've found the very similar issue: No callback get called from FSEventStreamCreate with modifications created by self in watched file. It seems the question didn't get any useful answer. Feeling like I'm missing something, but I don't know what. I've also changed latency argument to non zero value but got no luck.As long as this problem is not solved, I can't use FSEvents. So for now, I use
kqueue, which works perfectly fine for my purpose, on OS X...
2016-05-14
equal?の挙動
二週間前くらいにバグの報告を受けた(それ自体は修正済み)。バグの原因を突き詰めた際に「これはブログのネタになる」と思っていたのだが、それから随分時間が経ってしまった。多少賞味期限が切れてしまった気がしないでもないが、ちょっとした後方互換を壊す修正でもあるし、適当に記録に残しておく。ちなみにバグの報告はこれ。
バグの原因を要約すると以下の二つ:
1.に関してはレコード周りを取り除いてしまえば直るのは明白だったのだが、2.との兼ね合いでどうしようかなぁとい感じになるもの。そもそも、
SRFI-116はイミュータブルなリストを定義しているんだけど、その参照実装のテストケースが
じゃあどうしたか?実装をR6RSの
バグの原因を要約すると以下の二つ:
eqv?が循環構造のあるレコードを受け取るとSEGVを起こす。equal?がレコードの中身を比較する(R6RS的には規格違反)
eqv?がレコードの中身をチェックするのは、実はR6RS、R7RS両方で規格違反なのだが(明確にアドレス比較のみと書いてある)、equal?を変更した際にR6RSのテストスイートを通すために必要だったという経緯がある。(テストケースの意味もよく分からないんだけど、フィールドの存在しないレコードはコンストラクタが同一のオブジェクトを返してもいいってことなのかなぁ?)1.に関してはレコード周りを取り除いてしまえば直るのは明白だったのだが、2.との兼ね合いでどうしようかなぁとい感じになるもの。そもそも、
eqv?が中身を見るというのはいろいろおかしい感じがするので(リストの中身とか見ないわけだし)、取り除いてしまいたい感はあった。となると2.との兼ね合いだけなんだけど、これ自体はもともと便利だからという理由で入れてあっただけなので、正直取り除いてもそんなに影響ないだろうと思っていた。実際R7RS的には未規定なわけだし。が、SRFI-116に落とし穴が潜んでいた。SRFI-116はイミュータブルなリストを定義しているんだけど、その参照実装のテストケースが
equal?にレコードの中身を検査することを要求するコードになっていた。参照実装はChibiとChickenの2つの処理系で動くように実装されているのだが、どうもこの2つはレコードの中身を見るみたいである。新しいSRFIだからまだ使ってる人は少ないだろうし、暗黙的な要求なので無視しても問題ない気はしたのだが、人気処理系のうちの二つでやられているのなるとなぁという気持ちの方が大きかった。じゃあどうしたか?実装をR6RSの
equal?とそれ以外という風に分けた。ポータブルなコードを書くという上ではR6RSの方がより細かく規定してあるのでいいのだろうけど、これくらいの処理系拡張は許して欲しいという気持ちもある。そうすると実装を二つにする以外に方法が思いつかなかったのだ。ということで、(rnrs base)と(scheme base)で定義されているequal?は別物になった。この動作に依存したコードを書くというのはあんまりないと思うけど、ハッシュテーブルをequal?で作ってキーにレコードを使用している場合は影響がでるという話。
2016-05-08
肉体改造部 第十七週
ほぼ月一になっている気がする。
計量結果:
計量結果:
- 体重: 70.4kg (-0.5kg)
- 体脂肪率: 21.6% (-0.8%)
- 筋肉率:43.3% (+0.1%)
2016-04-20
言語レベル
(年に一回くらいこの手のことを書いてる気がしないでもないなぁ)
ある言語を話すことができるという理由で与えられるチャンスはそんなに多くないが、話せないから逃すチャンスというのは多々ある。これは英語に限ったことではなく、例えばここオランダでは募集要項にネイティブもしくはそれに準ずるオランダ語が話せること、ということが明確に書かれていることがある。逆に言うと、英語が話せればいいという職も多数ある。これが日本だと日本語は大前提になるのでExpatの多い国の特徴とも言えるのだろう。自分自身がどれくらい話せるのかとかを客観的に見たことがあまりないので、多少いろいろな角度からどの言語がどれくらいできるのか分析してみたくなった。
分析するにはある程度の基準がいる。とりあえず大きく5つのレベルに分けることにした。ただし、中間を表すために総数を20段階とし、5段階で区切るというようにする。例えば、日常会話はレベル5だが、ビジネスレベル(レベル10)に達していないが日常会話以上というのはレベル6から9の間といった具合である。以下はレベル:
さて、上記のレベルに自分の言語を当てはめてみる。何かしらのテストを受けて計測したというわけではないので、感覚的にという単なる目安である。
全体の習熟度とすればこんな感じなんだろうけど、個別にみると意外と面白いことが分かる。例えばくしゃみをした人にかける言葉として英語では「Bless you」、オランダ語では「Gezondheid」がある。オランダに長く住んでいるので誰かがくしゃみをすると、たとえくしゃみをした人がオランダ語を話せなくても、「Gezondheid」というようになった。同様なものに「Alstublieft」もしくは「Alsjeblieft」がある。もう少し込み入った例だと、「Kan ik pinnen?」ある。これは「Can I use debit?」のオランダ語バージョンと思ってもらえばいいのだが、アメリカに行ったときとか、店員さんがオランダ語喋れない場合でもこれが勝手に出てくる(アメリカで出た際は流石に、「Kan ik,,, can I use credit card?」になったが)。特定のシチュエーションに於いてはアウトプットが最も多いものが勝手に口をついてくるみたいである。
言語の習熟度があがると、言語間の壁のようなものが薄くなる気がする。最近は日本語のやたらいっぱい母音を喋らないといけないというのが面倒に感じるのだが、これが時として悪い方向に働く。例えば日本人と話しているのに、ふっと英語になるとか。これが起きるときは大体分かってて、
とりとめなく終わり。
ある言語を話すことができるという理由で与えられるチャンスはそんなに多くないが、話せないから逃すチャンスというのは多々ある。これは英語に限ったことではなく、例えばここオランダでは募集要項にネイティブもしくはそれに準ずるオランダ語が話せること、ということが明確に書かれていることがある。逆に言うと、英語が話せればいいという職も多数ある。これが日本だと日本語は大前提になるのでExpatの多い国の特徴とも言えるのだろう。自分自身がどれくらい話せるのかとかを客観的に見たことがあまりないので、多少いろいろな角度からどの言語がどれくらいできるのか分析してみたくなった。
分析するにはある程度の基準がいる。とりあえず大きく5つのレベルに分けることにした。ただし、中間を表すために総数を20段階とし、5段階で区切るというようにする。例えば、日常会話はレベル5だが、ビジネスレベル(レベル10)に達していないが日常会話以上というのはレベル6から9の間といった具合である。以下はレベル:
- レベル0:その言語を全く話せない
- レベル5:日常会話レベル、かなり大変だがその言語で生活できる
- レベル10:ビジネスレベル、職場でコミュニケーションができる
- レベル15:高等教育レベル、日本なら高校卒業時の国語
- レベル20:専門家レベル、この言語に関する知識で飯が食える
さて、上記のレベルに自分の言語を当てはめてみる。何かしらのテストを受けて計測したというわけではないので、感覚的にという単なる目安である。
- 日本語:レベル15(多分もう少し低いが、敬語とか忘れたし、日本の高校卒業してるということで)
- 英語:レベル13 - 14(多少色眼鏡付きな気もするけど)
- オランダ語:レベル4 - 6(一応生きていけるが辛い。仕事では使えない)
全体の習熟度とすればこんな感じなんだろうけど、個別にみると意外と面白いことが分かる。例えばくしゃみをした人にかける言葉として英語では「Bless you」、オランダ語では「Gezondheid」がある。オランダに長く住んでいるので誰かがくしゃみをすると、たとえくしゃみをした人がオランダ語を話せなくても、「Gezondheid」というようになった。同様なものに「Alstublieft」もしくは「Alsjeblieft」がある。もう少し込み入った例だと、「Kan ik pinnen?」ある。これは「Can I use debit?」のオランダ語バージョンと思ってもらえばいいのだが、アメリカに行ったときとか、店員さんがオランダ語喋れない場合でもこれが勝手に出てくる(アメリカで出た際は流石に、「Kan ik,,, can I use credit card?」になったが)。特定のシチュエーションに於いてはアウトプットが最も多いものが勝手に口をついてくるみたいである。
言語の習熟度があがると、言語間の壁のようなものが薄くなる気がする。最近は日本語のやたらいっぱい母音を喋らないといけないというのが面倒に感じるのだが、これが時として悪い方向に働く。例えば日本人と話しているのに、ふっと英語になるとか。これが起きるときは大体分かってて、
- グループ内に日本人以外がいる
- カタカナ語を英語の発音で喋ってしまう
とりとめなく終わり。
2016-04-19
Inter-operable hidden binding
The subject might sound weird, but I couldn't find any better name. So please bare with it.
Problem
Suppose you have a situation that 2 macros need to refer the same implicitly bound variable. For example, considerdefine-generic and define-method; the declaration of generic function is done by define-generic, and adding definition to it is done by define-method. Now, you want to write it as simple as possible, so you've decided to use implicitly bound hashtable.
;; naive definition of define-generic (doesn't work)
(define-syntax define-generic-naive
(syntax-rules ()
((_ name)
(begin
(define implicit (make-eq-hashtable))
(define (name . args)
;; lookup and execute
)))))
(define-syntax define-method-naive
(syntax-rules ()
((_ name formals body ...)
(begin
(define (real-proc . formals) body ...)
(define dummy
;; oops, implicit can't be referred here!
(hashtable-set! implicit 'name real-proc))))))
Now, your task is make this happen somewhat.Passing explicitly
Taking this path isn't really what I want, but it's the only way to do it on R7RS.
(define-syntax define-generic-explicit
(syntax-rules ()
((_ name table)
(begin
(define table (make-eq-hashtable))
(define (name . args)
;; lookup and execute
)))))
(define-syntax define-method-naive
(syntax-rules ()
((_ name table formals body ...)
(begin
(define (real-proc . formals) body ...)
(define dummy
(hashtable-set! table 'name real-proc))))))
The problem with this implementation is that you need to know the name of the shared bindings. It might be good for debugging or breaking, but you probably don't want to care about something only used internally.Macro generating macro
If 2 macros cannot refer the variable defined in one of the macro, then make it in the one macro like this:(define-syntax define-generic/defmethod
(syntax-rules ()
((_ name method-name)
(begin
(define shared (make-eq-hashtable))
(define (name . args)
;; lookup and execute
)
(define-syntax method-name
(syntax-rules ()
((_ name shared formals body (... ...))
(begin
(define (real-proc . formals) body (... ...))
(define dummy
(hashtable-set! shared 'name real-proc))))))))))
It's probably better than explicitly passing, but it's rather ugly. The method definition should be more generic. In this implementation, the method definition belongs to specific generic function definition.Identifier macro
If you are using R6RS, thensyntax-case can handle non list macro (not sure how it should be called, but I say identifier macro). So if the name of generic function itself can be evaluated to implicit definition name, then we can share the binding by referring the name.
(define-syntax define-generic
(syntax-rules ()
((_ name shared)
(begin
(define shared (make-eq-hashtable))
(define (real-proc . args)
;; lookup
)
(define-syntax name
(lambda (x)
(syntax-case x ()
((_ args (... ...)) #'(real-proc args (... ...)))
(k (identifier? #'k) #'shared))))))))
(define-syntax define-method
(syntax-rules ()
((_ name formals body ...)
(begin
(define (real . formals) body ...)
;; method name should generic function name;
;; thus, it's an identifier macro to return
;; implicit table name.
(define dummy (hashtable-set! name 'name real))))))
Better, at least for me. If I see it with half eye closed, then it looks like fake LISP-2.Pitfalls I've got
I first thought that maybe I can usedatum->syntax to create the same name; however, this wasn't a good idea. It is okay to if both generic function declaration and method definitions are in the same library; otherwise, you'd get a problem. Suppose you have a library (a) contains only define-generic and other library (b) contains define-method. Now, which template identifier you should use to generate the same name of the implicit binding? You need to use the identifier define-generic in the library (a), and it's impossible to use it in library (b). (This is the reason why I needed to write the version 2, macro generating macro.)Conclusion
I don't have any intention to say, macro is the best, or something like that, but if something beyond procedure (in this case, emulating LISP-2, kind of), then it is rather necessary feature.2016-04-15
Generic record copy
I've found a tweet that says R7RS
Now, if I just say like this, then it's not so fun. So let's write kind of generic copy procedure. Before that, here our definition of copy is deep copy. So it creates a new object without having the same object inside. So more like cloning.
NB:
Then you can use it like this:
The whole scripts are here:
define-record-type doesn't create copier (or copy constructor) by default. Well, I can imagine why it doesn't if I think of C++'s copy constructor (which I think very confusing and causing unexpected behaviour). And it's also context dependent what exactly copy
means.Now, if I just say like this, then it's not so fun. So let's write kind of generic copy procedure. Before that, here our definition of copy is deep copy. So it creates a new object without having the same object inside. So more like cloning.
Preparation
If it's generic, then it should work also builtin types. Generally, Scheme chose to have distinct procedures per types and what we want is one generic procedure. The very simple strategy would be dispatching. It might be convenient if users can specify how copy works per types. So the interface of copy procedure would look like this:(define *copier-table* '())
(define (generic-copy obj)
(cond ((assoc obj *copier-table* (lambda (x p) (p x))) =>
(lambda (s) ((cdr s) obj)))
;; shallow copy, sort of
(else obj)))
(define (register-copier! pred copier)
(set! *copier-table* (cons (cons pred copier) *copier-table*)))
To register built-in types, we can do like this:
(register-copier! pair? list-copy) (register-copier! vector? vector-copy) (register-copier! string? string-copy) (register-copier! bytevector? bytevector-copy)Now, we have generic copy procedure for built-in types.
NB:
list-copy and vector-copy doesn't consider the elements of copying object. If you want to follow the definition of copy here, you need to create own copy procedure.Syntax
You know howdefine-record-type works, right? It needs to be fed name of constructors, predicate procedures. So doing the same for copy procedure. Let's call our brand new record definition syntax define-record-type/copy. It would look like this:
(define-record-type/copy pare (kons a d) pare? pare-copy (a kar) (d kdr))The extra argument
pare-copy is the procedure automatically generated by the macro.Implementation strategy
Now, how can we implement it? The strategy I chose (and probably this is the only way to do it portably) is that:- Collect field value and order it by constructor tag
- Create object by passing above value with specified constructor
- Set field values of fields which are not listed on constructor
(define-syntax define-record-copier
(syntax-rules ()
((define-record-copier "emit" name (ctr f ...) (acc ...) ((a m) ...))
;; now we have all information
(define (name obj)
(let ((c (ctr (acc obj) ...)))
;; mutate if mutators are defined, then we use it.
;; to make it simple, we do for all mutator. so some
;; of them are just useless.
;; FIXME this is not efficient.
(m c (a obj)) ...
c)))
((_ "mutator" name ctr accessor mutator ())
(define-record-copier "emit" name ctr accessor mutator))
((_ "mutator" name ctr accessor (mutator* ...) ((f a) rest ...))
(define-record-copier "mutator" name ctr accessor
(mutator* ...) (rest ...)))
((_ "mutator" name ctr accessor (mutator* ...) ((f a m) rest ...))
(define-record-copier "mutator" name ctr accessor
(mutator* ... (a m)) (rest ...)))
((_ "collect" name ctr (acc ...) () (def* ...))
(define-record-copier "mutator" name ctr (acc ...) ()(def* ...)))
((_ "collect" name ctr (acc ...) (field field* ...) (def* ...))
(begin
;; this part is not R7RS portable since 'foo' doesn't have to be
;; renamed (right?). so some of implementation may raise an error
;; of redefinition (e.g. Foment)
;; however we can't use letrec-syntax because it creates a scope.
;; sucks...
(define-syntax foo
(syntax-rules (field)
((_ ?n ?c
((field ac . ignore) rest (... ...))
(next (... ...))
(src (... ...)))
(define-record-copier "collect" ?n ?c (acc ... ac)
(next (... ...)) (src (... ...))))
((_ ?n ?c (_ rest (... ...)) (next (... ...)) (src (... ...)))
(foo ?n ?c (rest (... ...)) (next (... ...)) (src (... ...))))))
(foo name ctr (def* ...) (field* ...) (def* ...))))
((_ name ctr (ctr-field* ...) (field-def* ...))
(define-record-copier "collect" name ctr
() ;; accessor
(ctr-field* ...)
(field-def* ...)))))
(define-syntax define-record-type/copy
(syntax-rules ()
((_ name (ctr field* ...) pred copier field-def* ...)
(begin
(define-record-type name (ctr field* ...) pred
field-def* ...)
(define-record-copier copier (ctr field* ...)
(field* ...) (field-def* ...))))))
I usually use letrec-syntax to detect free identifier (well, it should be bound identifier but I don't think there's no way to do it in range of R7RS). But needed to use define-syntax (see comment).Then you can use it like this:
(define-record-type/copy pare (kons a d) pare? pare-copy
(a kar)
(d kdr)
(s pare-src pare-src-set!))
(register-copier! pare? pare-copy)
(let ((p (kons 'a 'b)))
(pare-src-set! p '(src))
(let ((c (generic-copy p)))
(print (kar c))
(print (kdr c))
(print (pare-src c))))
(Write your own print procedure :P). The implementation is not efficient since we call mutator procedure no matter what. To make it efficient, you need to get mutators of which are not listed on constructor tags. The whole scripts are here:
Conclusion
Use R6RS or SRFI-99.2016-04-11
肉体改造部 第十四週
先週は何故か忘れた。
計量結果:
計量結果:
- 体重: 70.9kg (-0.3kg)
- 体脂肪率: 22.4% (+0.3%)
- 筋肉率:43.2% (±0.0%)
2016-04-08
mod-exptの高速化
タイトルは大分嘘です。
Linux上での暗号ライブラリテストが以上に遅かった。他のOSでは問題ないのだが、Linuxだけ10倍以上遅い。何が遅いのかなぁと調べてみると、鍵対の生成が1024ビット程度でも3秒くらいかかっているというものだった。これはおかしいなぁと思っておもむろに鍵対生成をプロファイラにかけてみると
この手続き自体は確かに重いものなのだが、どうもおかしい。以前(多分0.5.x辺り)ではそんなに時間がかかった記憶がない。つまりその辺から今までで入れた変更でLinuxのみが遅くなった可能性がある。記憶を辿ってみると確かにBignumの演算に手を入れた記憶があったので、とりあえずソースを覗いてみる。っが、特に不審な部分も見当たらない。Linux固有の何かを使ったものはないという意味でではあるが。
疑問を疑問のままにしておくのは今一気持ちが悪いので、Valgrindについてるcallgrindを使ってCレベルのプロファイルを取る。すると、
具体的にはスタックベースを取得するのが異常に遅いっぽかった。そもそもスタックベースなど一回取得してしまえば変更されることはないはずなので毎回値を律儀に取得しにいくこともないよなぁと思いスレッドローカルな静的領域に格納するように変更。これだけで100倍の高速化に成功した。(実際は100倍の低速化が行われているので、元に戻っただけだが・・・)
ここからは(も?)与太話。
スタックベースの取得にはBoehmGCの
っで、Cygwinの実装を見てみた。
特に何もなく、callgrindが便利だったというだけの話だったりはする。他のプロファイラと違いランタイムにリンクさせる必要ないというのはとてもありがたい。その分処理は劇的に遅くなるけど、的が絞れているならこれほど便利なものはないなぁと思ったのでした。
Linux上での暗号ライブラリテストが以上に遅かった。他のOSでは問題ないのだが、Linuxだけ10倍以上遅い。何が遅いのかなぁと調べてみると、鍵対の生成が1024ビット程度でも3秒くらいかかっているというものだった。これはおかしいなぁと思っておもむろに鍵対生成をプロファイラにかけてみると
mod-exptが遅い。120回程度呼ばれて1500ms消費という感じであった。この手続き自体は確かに重いものなのだが、どうもおかしい。以前(多分0.5.x辺り)ではそんなに時間がかかった記憶がない。つまりその辺から今までで入れた変更でLinuxのみが遅くなった可能性がある。記憶を辿ってみると確かにBignumの演算に手を入れた記憶があったので、とりあえずソースを覗いてみる。っが、特に不審な部分も見当たらない。Linux固有の何かを使ったものはないという意味でではあるが。
疑問を疑問のままにしておくのは今一気持ちが悪いので、Valgrindについてるcallgrindを使ってCレベルのプロファイルを取る。すると、
mod-exptの処理自体は高速に終わっているという結果が取れた。っで、コールグラフのその下を見ると、スタック領域の割り出しの処理が異常に重たい。そういえば、Bignumの計算でスタックが溢れる不具合を直した際にそんな処理入れたなぁと思い、ダミーの値を返すようにしてSchemeのプロファイルを取る。3秒が30msになった。お前か・・・具体的にはスタックベースを取得するのが異常に遅いっぽかった。そもそもスタックベースなど一回取得してしまえば変更されることはないはずなので毎回値を律儀に取得しにいくこともないよなぁと思いスレッドローカルな静的領域に格納するように変更。これだけで100倍の高速化に成功した。(実際は100倍の低速化が行われているので、元に戻っただけだが・・・)
ここからは(も?)与太話。
スタックベースの取得にはBoehmGCの
GC_get_stack_baseを使っているのだが、LinuxとCygwinで100倍以上の差が付くのはなぜだろうと思いちょっと実装を覗いてみた。Linux(x86_64)では以下の処理を行う:pthread_getattr_npの呼び出しpthread_attr_getstackの呼び出しpthread_attr_destroyの呼び出し
っで、Cygwinの実装を見てみた。
GC_API int GC_CALL GC_get_stack_base(struct GC_stack_base *sb)
{
void * _tlsbase;
__asm__ ("movl %%fs:4, %0"
: "=r" (_tlsbase));
sb -> mem_base = _tlsbase;
return GC_SUCCESS;
}
以上!そら速いわ。。。1万回呼び出されても誤差の範囲に収まるだろうなぁというのは想像に難くない。特に何もなく、callgrindが便利だったというだけの話だったりはする。他のプロファイラと違いランタイムにリンクさせる必要ないというのはとてもありがたい。その分処理は劇的に遅くなるけど、的が絞れているならこれほど便利なものはないなぁと思ったのでした。
2016-04-01
プロセスとI/O
サーバーが正常に立ち上がったかどうかを確認するのに起動ログを見るか、実際にアクセスして動かないことを確認するしかないというのがだるくなった。なので、ファイルを監視しつつ失敗のキーワードがあれば通知するものを作ったのだが、どうもプロセスをデタッチすると何も出力されないことに気付いた。Sagittarius 0.7.2までは子プロセスの標準入出力は常にパイプが割り当てられるのだが、親プロセスが終了するとパイプから出力を読み取るプロセスがなくなるので何も出力されないという話だった。これでは不便だなぁと思い、えいや!っと出力先を制御できるようにしてみた。
こんな感じで使う。
これ変更したのはいいけど、実際に通知を行うのにコンソールに垂れ流すと見落とすということで、
こんな感じで使う。
(import (rnrs) (sagittarius process))
(let ((proc (make-process "foo" '("process")))
(outfile "pout"))
(process-call proc :output outfile))
これで、プロセスの標準出力はpoutというファイルになる。出力先を標準出力にしたいときは:stdoutを使う。もちろん、入力(:inputキーワード引数)とエラー出力(:errorキーワード引数)もサポートしている。便利手続きのcreate-processもこれを考慮するように変更したいが、まだしてない(こっちはパイプでも問題ないようにしか使ってないとも言う)。今のところ出力先ファイルは上書きででしか開けないが、必要があれば追記できるようにするかもしれない。これ変更したのはいいけど、実際に通知を行うのにコンソールに垂れ流すと見落とすということで、
notify-sendコマンド使ってデスクトップに通知するようにしたら、パイプ使ってても問題ないなくなった。変更自体は有用だと思うけど、最優先で変更したのはいいが必要なくなった子になってしまった。
2016-03-29
Ellipses expansion of syntax-rules
An interesting post was posted on c.l.s. (c.f. Nested ellipses) It's about how ellipses of syntax-rules should be expanded. The code is as follows:
This, I believe, restricts that input expressions of the multiple ellipses have the same length of input. In above example, x and y should have the same length. R7RS, on the other hand, requires to consume all input. Thus, x and y may have different length of inputs (e.g.
Maybe there's more direct statement which specifies the behaviour of this case.
(define-syntax test
(syntax-rules ()
((test (x ...) ((y ...) ...) )
'((x (x y) ...) ...) ) ) )
(test (a b c)
((1 2 3) (4 5 6) (7 8 9)) )
I'm not sure if this is an error since (x (x y) ...) contains 2 times x followed by an ellipsis. (I think it is, and SRFI-72 expander, a.k.a Van Tonder expander, signals an error.) So I've removed the first x and tested on couple of R6RS and R7RS implementations.
;; Removed the first x
(define-syntax test
(syntax-rules ()
((test (x ...) ((y ...) ...) )
'(((x y) ...) ...) ) ) )
(test (a b c)
((1 2 3) (4 5 6) (7 8 9)) )
#|
Either:
#1
(((a 1) (b 2) (c 3))
((a 4) (b 5) (c 6))
((a 7) (b 8) (c 9)))
Or
#2
(((a 1) (a 2) (a 3))
((b 4) (b 5) (b 6))
((c 7) (c 8) (c 9)))
|#
Implementations emit the #1 are the following:- All R6RS implementations
- Foment
- Chibi
- Sagittarius using
(scheme base)library - Gauche
- Picrin
Pattern variables that occur in subpatterns followed by one or more ellipses may occur only in subtemplates that are followed by (at least) as many ellipses. These pattern variables are replaced in the output by the input subforms to which they are bound, distributed as specified.
R6RS: 11.19 - Macro transformers
Pattern variables that occur in subpatterns followed by one or more instances of the identifier ellipsis are allowed only in subtemplates that are followed by as many instances of ellipsis . They are replaced in the output by all of the elements they match in the input, distributed as indicated.I think the difference between R6RS and R7RS is matched ellipses consuming part. R6RS also requires the following:
R7RS: 4.3.2 - Pattern language
The subtemplate must contain at least one pattern variable from a subpattern followed by an ellipsis, and for at least one such pattern variable, the subtemplate must be followed by exactly as many ellipses as the subpattern in which the pattern variable appears. (Otherwise, the expander would not be able to determine how many times the subform should be repeated in the output.)
This, I believe, restricts that input expressions of the multiple ellipses have the same length of input. In above example, x and y should have the same length. R7RS, on the other hand, requires to consume all input. Thus, x and y may have different length of inputs (e.g.
(test (a b c) ((1 2 3 10) (4 5 6) (7 8 9))) should be valid on above example). In such cases, the expander can't determine how it should be expanded if it needs to expand like R6RS does (as R6RS mentioned).Maybe there's more direct statement which specifies the behaviour of this case.
2016-03-27
肉体改造部 第十二週
風邪ひいて一週間マルッと寝込んでいたりして二週飛ばし。こんなんばっかだな。。。
計量結果:
なんだかんだで食べ過ぎるなぁという感じがあるので、頑張って食欲に負けないようにしないといけないのだが、「ダイエットは明日から」という言葉を使う人の気持ちが分かるレベルで誘惑に負けそうになる(というか負けてる)。
計量結果:
- 体重: 71,2kg (+1.7kg)
- 体脂肪率: 22.1% (+1.6%)
- 筋肉率:43.2% (+0.6%)
なんだかんだで食べ過ぎるなぁという感じがあるので、頑張って食欲に負けないようにしないといけないのだが、「ダイエットは明日から」という言葉を使う人の気持ちが分かるレベルで誘惑に負けそうになる(というか負けてる)。
2016-03-26
ファイルシステムの監視 実装編
とりあえず、inotify、kqueueとReadDirectoryChangesWの3つで大体同じように動くものができた。例えばtailコマンドっぽい何かは以下のように書くことができる。
実装に関して
前回も書いたが三者三様なのでそれぞれ苦労した。inotifyは素直にできているのでLinuxのinotify(7)にあるサンプルを参考にしながらで十分だった。思ったとおりこれが一番楽だった。次いでkqueueなのだが、こいつは例があまりなかったのと、kqueue自体が非常に総称的にできているので理解するまでに苦労した。理解してしまえばまぁそれほどという感じではある。ReadDirectoryChangesWはこれ自体はそんなに複雑じゃないんだけど、どちらかというとそれ以外の部分(OVERLAPPEDとかFILE_NOTIFY_INFORMATIONとか)が多少面倒だった感じ。動いてるけど正しく実装したのか自信ない。
実装間に於ける制限
意外だったのはkqueueが制限が一番大きくなったこと。kqueueはファイルの監視はできるけど、ディレクトリの監視をした際にどのファイルが変更されたとか追加されたとかを知る術がない。ものすごく頑張ればやれなくないんだけど、監視対象毎にファイルディスクリプタが必要になるので、万を超えるファイルとかがあるディレクトリの監視とかすると普通に死にそう。(ものすごく頑張る必要性を今のところ感じていないので頑張っていないが・・・)
次いでinotify。ディレクトリの再帰的監視は頑張らないと無理(kqueueも無理だけど)。まぁ、再帰的に監視したいかと言われるとよく分からないが。とりあえず他の実装もこれにあわせるようにしてお茶を濁した。
Windowsはファイルの監視ができないんだけど、ディレクトリの監視をすればどのファイルが変更されたかの情報が取れるので特に問題なかった。ありそうなのは、Windows Vista以降ではデフォルトでアクセス時間の変更がされないので、それの監視をしたい場合はシステムを弄らないといけないことか(実装とは関係ない) 。
落とし穴
Cygwinが実は落とし穴だった。CygwinはPOSIX環境を提供してくれるんだけど、inotifyもkqueueもPOSIXじゃないので存在しない。そしてCygwin自体にファイル等を監視するようなAPIもない。どうしたかといえば、Windowsの実装の上にパス変換(
所感
疲れた。後は使いつついじっていく感じかな。
;; tail.scm
(import (rnrs) (getopt) (sagittarius filewatch) (binary io))
(define (tail file offset)
(define watcher (make-filesystem-watcher))
(define in (open-file-input-port file))
;; dump contents to stdout
(define (dump)
(let loop ()
(let ((line (get-line in)))
(unless (eof-object? line)
(put-bytevector (standard-output-port) line)
(put-bytevector (standard-output-port) #vu8(10))
(loop)))))
(define size (file-size-in-bytes file))
;; move port position if the size if more than offset
(when (> size offset) (set-port-position! in (- size offset)))
;; dump first
(dump)
;; add path to file watcher
(filesystem-watcher-add-path! watcher file '(modify)
(lambda (path event) (dump)))
;; monitor on foreground.
(filesystem-watcher-start-monitoring! watcher :background #f))
;; this tail is not line oriented
;; it shows tail of the file from the given offset.
(define (main args)
(with-args (cdr args)
((offset (#\o "offset") #t "1024")
. rest)
(tail (car rest) (string->number offset))))
#|
sash tail.scm foo
|#
これで延々とファイルに追加されたものを標準出力に吐き出していく。まだドキュメント化していないが、捻りを加える必要もないだろうし、多分これが最終形になると思われる。実装に関して
前回も書いたが三者三様なのでそれぞれ苦労した。inotifyは素直にできているのでLinuxのinotify(7)にあるサンプルを参考にしながらで十分だった。思ったとおりこれが一番楽だった。次いでkqueueなのだが、こいつは例があまりなかったのと、kqueue自体が非常に総称的にできているので理解するまでに苦労した。理解してしまえばまぁそれほどという感じではある。ReadDirectoryChangesWはこれ自体はそんなに複雑じゃないんだけど、どちらかというとそれ以外の部分(OVERLAPPEDとかFILE_NOTIFY_INFORMATIONとか)が多少面倒だった感じ。動いてるけど正しく実装したのか自信ない。
実装間に於ける制限
意外だったのはkqueueが制限が一番大きくなったこと。kqueueはファイルの監視はできるけど、ディレクトリの監視をした際にどのファイルが変更されたとか追加されたとかを知る術がない。ものすごく頑張ればやれなくないんだけど、監視対象毎にファイルディスクリプタが必要になるので、万を超えるファイルとかがあるディレクトリの監視とかすると普通に死にそう。(ものすごく頑張る必要性を今のところ感じていないので頑張っていないが・・・)
次いでinotify。ディレクトリの再帰的監視は頑張らないと無理(kqueueも無理だけど)。まぁ、再帰的に監視したいかと言われるとよく分からないが。とりあえず他の実装もこれにあわせるようにしてお茶を濁した。
Windowsはファイルの監視ができないんだけど、ディレクトリの監視をすればどのファイルが変更されたかの情報が取れるので特に問題なかった。ありそうなのは、Windows Vista以降ではデフォルトでアクセス時間の変更がされないので、それの監視をしたい場合はシステムを弄らないといけないことか(実装とは関係ない) 。
落とし穴
Cygwinが実は落とし穴だった。CygwinはPOSIX環境を提供してくれるんだけど、inotifyもkqueueもPOSIXじゃないので存在しない。そしてCygwin自体にファイル等を監視するようなAPIもない。どうしたかといえば、Windowsの実装の上にパス変換(
cygwin_conv_path)を噛ませるようにした。また、監視を止めるのに他の環境だとスレッドに割り込みをかけるようにしているが、Windowsのコードを流用しなければならないのでそれができない(Cygwinはスレッドの割り込みにシグナルを使うがWindowsはSetEventを使っている)。しょうがないので泥臭い方法で回避している。所感
疲れた。後は使いつついじっていく感じかな。
2016-03-22
ファイルシステムの監視
最近ファイルの変更を検知したいという要望が僕の中であがってきている。追記型のログファイルを調べるとかそんなちょっとしたことからなんだけど、あると便利かなぁと思い始めてきた。っで、いろいろ調べてみた結果OS毎に作法が全然違うという悲しい事実が判明。Linuxはinotify、WindowsはReadDirectoryChangesW、*BSDはkqueue、OSXはFSEvents(だけど、kqueueも使えるっぽいのでそっち使う予定。手元にOSないし)みたいである。
ちらっと調べてみた感じでは、それぞれに一長一短あるんだけど、個人的にはinotifyが一番楽っぽいイメージ。次いでkqueue。Windowsはやりたいことはやれなくないけど結構大変っぽい。Windowsの問題はファイルシステムの監視がフォルダのみというところで、ファイル自体の変更を検知する直接的な方法はないっぽい。パスを分解、フォルダを監視、その上でターゲットのファイルが変更されたかどうかをチェックするという方法になりそう。
一番楽っぽいかなと思われるLinuxの実装は既にリポジトリに入れた。どのイベントを取るかとか、ディレクトリが指定された際はどうするかとかまだ考えないといけないけど、単純なファイルの監視というところは動いている。次はWindows+Cygwinを何とかしつつ、kqueueは最後にやる感じ。一番問題になるのは、動作をそろえるところだろうなぁ。こればっかりは地道にやるしかないので、ある程度作ったら使いながら調節するという感じになるような気がする。
ちらっと調べてみた感じでは、それぞれに一長一短あるんだけど、個人的にはinotifyが一番楽っぽいイメージ。次いでkqueue。Windowsはやりたいことはやれなくないけど結構大変っぽい。Windowsの問題はファイルシステムの監視がフォルダのみというところで、ファイル自体の変更を検知する直接的な方法はないっぽい。パスを分解、フォルダを監視、その上でターゲットのファイルが変更されたかどうかをチェックするという方法になりそう。
一番楽っぽいかなと思われるLinuxの実装は既にリポジトリに入れた。どのイベントを取るかとか、ディレクトリが指定された際はどうするかとかまだ考えないといけないけど、単純なファイルの監視というところは動いている。次はWindows+Cygwinを何とかしつつ、kqueueは最後にやる感じ。一番問題になるのは、動作をそろえるところだろうなぁ。こればっかりは地道にやるしかないので、ある程度作ったら使いながら調節するという感じになるような気がする。
2016-03-19
ズンドコキヨシ
風邪(と思われる)で1週間寝込んでいたのだが、体調がほぼ戻ってきた。Twitter等でズンドコキヨシなるものを見かけたので、1週間ぶりにリハビリを兼ねて書いてみた。(1週間もコード書かないと鈍るよね?)
#!r6rs
(import (rnrs) (srfi :27))
(define zun "ズン")
(define doko "ドコ")
(define kiyoshi "キ・ヨ・シ!")
(define (zun? o) (string=? zun o))
(define (doko? o) (string=? doko o))
(define (kiyoshi! o) (display o) (display kiyoshi) (exit 0))
(define zun&doko (vector zun doko))
(define (zundoko-generator) (vector-ref zun&doko (random-integer 2)))
(define init-state 0)
(define (gen-next n) (lambda (o) (display o) n))
(define ->init (gen-next init-state))
(define states
`#((,zun? ,(gen-next 1) ,->init)
(,zun? ,(gen-next 2) ,->init)
(,zun? ,(gen-next 3) ,->init)
(,zun? ,(gen-next 4) ,->init)
(,doko? ,kiyoshi! ,(gen-next 4)) ;; more than 4 zun, loop it
))
(random-source-randomize! default-random-source)
(let loop ((ns init-state))
(let ((token (zundoko-generator))
(state (vector-ref states ns)))
(if ((car state) token)
(loop ((cadr state) token))
(loop ((caddr state) token)))))
普通に4回と1回を数えた方がすっきりするような気もしないでもない。
2016-03-07
オランダの転職エージェント
Amazonのオファーを蹴ってしまった+収入を増やしたい(いろいろ要りようなのですよ)という思いから転職活動をしている。なんだかんだ(いろいろリスクはあるが)で手っ取り早く収入を増やすならよりよい条件の職場に行くのが早い。そうはいっても、あまりがつがつ転職活動をするというつもりもなく、いい条件の職があったらという感じでゆる~くやっているのではあるが。
オランダに限らず転職は3種類くらいパターンがあると思う。
さて、これだけならわざわざブログの記事にすることもないのだが、ちょっと頭にきたことがあったりして吐き出しを兼ねて適当にエージェントを使うのが嫌いな理由を書いていく。
【転職理由】
大抵のエージェントは、なんでこの会社にいきたいか、みたいな事を質問してくる。ついでになんで転職するのかとか。一度素直に「金」と答えたら、「それじゃだめだ」みたいなことを言われたことがある。仕事内容なんてお前らが言ってることと一致したことねえんだよ!転職理由なんて金払いがよければそれで十分だ!僕にとって「やりがい」とかは副次的であって、主目的は「金」だよ!「やりがい」で腹は膨れないっつーの!
【希望収入】
少なくともオランダでこれを聞かれるときには、セットで現在の収入も答える必要がある。ここで、現在の収入を真面目に答えると損をするので必ず月収なら500ユーロは多く答えておくとよい(経験談)。
転職するのであれば、よほど今の会社から逃げたいとかを除いて、収入が上がることを期待したいものである。よくも悪くも転職にはリスクが伴うし(ペンションとか、職歴とか)。っで、現在の収入より月500ユーロ多くもらえるのを希望すると、「500ユーロも増やせると思うの?」とか「現在の収入からみて妥当なところを探す」とか言われることがある(3分の2のの確立)。ぶっちゃけ、これを言われたら萎える(萎えた、今日)。Nettで300ユーロ増やすのがそんなに悪ですか?あの手この手使って最低ラインを下げようとしてますよね?ぶっちゃっけ月100ユーロ増える程度では職変えないぜ、普通。お前らの交渉能力の低さを棚に上げてこっちにばかり妥協点を押し付けんな!あんまり声を荒げるとか、大人気なく喧嘩する気もないので、希望額に届かなかったら容赦なく辞退するだけですよ。
【自己矛盾】
今の会社もエージェント経由だったんだけど、どうも同じ会社っぽいんだよね。当時(一年前だが)の担当(今の担当の上司らしい)は、この会社に採用された人は長く続けてるからこの会社はいい会社だ、みたいなこと言ってたんだけどねぇ。まぁ、人売り人買いの会社なんて早々に転職させて金儲けしてるんだから当然なのかもしれないけど、この節操のなさにはこっちもびっくりですよ。
【恩着せがましい】
「この会社に他の人を送るのストップしてる」というのは彼らの殺し文句である。そっちの事情は知ったことではないのだよ。それで恩を売って、決まった際に「これだけやったんだから給料が低くても転職しろ」みたいな態度に出られてはたまらない。お前らは仕事、こっちはリスクを負う、恩も義理もない。ビジネスでやってるのに、人情を人質に取ろうとするのに反吐がでる。多少の害には目をつぶれってか?ふざけんな!
適当に書きなぐってしまった。使えそうなら使うくらいの立場でいた方がいいということ。変に義理立てしたり、向こうの意味不明な言論に左右されないというのが大事である。あぁ、腹立った。
オランダに限らず転職は3種類くらいパターンがあると思う。
- 自分で探す
- 向こうから声をかけられる
- エージェント経由
- エージェントから連絡がくる
- 先方との面接日時を決める
- 面接
- エージェント経由でフィードバックを受け取る
- 採用、不採用が決まるまで2-3までを繰り返す
さて、これだけならわざわざブログの記事にすることもないのだが、ちょっと頭にきたことがあったりして吐き出しを兼ねて適当にエージェントを使うのが嫌いな理由を書いていく。
【転職理由】
大抵のエージェントは、なんでこの会社にいきたいか、みたいな事を質問してくる。ついでになんで転職するのかとか。一度素直に「金」と答えたら、「それじゃだめだ」みたいなことを言われたことがある。仕事内容なんてお前らが言ってることと一致したことねえんだよ!転職理由なんて金払いがよければそれで十分だ!僕にとって「やりがい」とかは副次的であって、主目的は「金」だよ!「やりがい」で腹は膨れないっつーの!
【希望収入】
少なくともオランダでこれを聞かれるときには、セットで現在の収入も答える必要がある。ここで、現在の収入を真面目に答えると損をするので必ず月収なら500ユーロは多く答えておくとよい(経験談)。
転職するのであれば、よほど今の会社から逃げたいとかを除いて、収入が上がることを期待したいものである。よくも悪くも転職にはリスクが伴うし(ペンションとか、職歴とか)。っで、現在の収入より月500ユーロ多くもらえるのを希望すると、「500ユーロも増やせると思うの?」とか「現在の収入からみて妥当なところを探す」とか言われることがある(3分の2のの確立)。ぶっちゃけ、これを言われたら萎える(萎えた、今日)。Nettで300ユーロ増やすのがそんなに悪ですか?あの手この手使って最低ラインを下げようとしてますよね?ぶっちゃっけ月100ユーロ増える程度では職変えないぜ、普通。お前らの交渉能力の低さを棚に上げてこっちにばかり妥協点を押し付けんな!あんまり声を荒げるとか、大人気なく喧嘩する気もないので、希望額に届かなかったら容赦なく辞退するだけですよ。
【自己矛盾】
今の会社もエージェント経由だったんだけど、どうも同じ会社っぽいんだよね。当時(一年前だが)の担当(今の担当の上司らしい)は、この会社に採用された人は長く続けてるからこの会社はいい会社だ、みたいなこと言ってたんだけどねぇ。まぁ、人売り人買いの会社なんて早々に転職させて金儲けしてるんだから当然なのかもしれないけど、この節操のなさにはこっちもびっくりですよ。
【恩着せがましい】
「この会社に他の人を送るのストップしてる」というのは彼らの殺し文句である。そっちの事情は知ったことではないのだよ。それで恩を売って、決まった際に「これだけやったんだから給料が低くても転職しろ」みたいな態度に出られてはたまらない。お前らは仕事、こっちはリスクを負う、恩も義理もない。ビジネスでやってるのに、人情を人質に取ろうとするのに反吐がでる。多少の害には目をつぶれってか?ふざけんな!
適当に書きなぐってしまった。使えそうなら使うくらいの立場でいた方がいいということ。変に義理立てしたり、向こうの意味不明な言論に左右されないというのが大事である。あぁ、腹立った。
2016-03-06
肉体改造部 第九週
なんかいろいろあって2週ほど飛ばしてしまった。
計量結果:
最近懸垂が普通に10回x4セットくらいできるようになってきたのでちょいちょい筋肉が付いてきたのではと思っているのだが、数字には表れていない様子(当てになるかもよく分からんけど)。懸垂しても上腕二頭筋にあまり負荷がかかっていない感じがするということは、自重トレーニングする分には十分ということなのだろうか?ジムに行きたいところではあるが、時間が取れないんだよなぁ。
計量結果:
- 体重: 72.9kg (+0.1kg)
- 体脂肪率: 23.7% (+0.1%)
- 筋肉率:42.6% (±0.0%)
最近懸垂が普通に10回x4セットくらいできるようになってきたのでちょいちょい筋肉が付いてきたのではと思っているのだが、数字には表れていない様子(当てになるかもよく分からんけど)。懸垂しても上腕二頭筋にあまり負荷がかかっていない感じがするということは、自重トレーニングする分には十分ということなのだろうか?ジムに行きたいところではあるが、時間が取れないんだよなぁ。
2016-03-04
Cache
I'm currently working on ORM library (this) and have figured out that creating prepared statement is more expensive than I expected. You might be curious how much more expensive? Here is the simple benchmark script I've used.
Now, my ORM library hides low level operations such as creating prepared statement, connection management, etc. (that's what ORM should do, isn't it?). So keeping prepared statement in users' script isn't an option. Especially, there's no guarantee that users woudl get the same connection each time they do some operation. So it's better to manage it on the framework.
Sagittarius has
There are variety of cache algorithms. Implementing all of them would take a bit time. So it's better to make a framework or interface of cache. The framework/interface should have the following properties:
(import (rnrs)
(time)
(sagittarius control)
(postgresql))
(define conn (make-postgresql-connection
"localhost" "5432" #f "postgres" "postgres"))
;; prepare the environment
(postgresql-open-connection! conn)
(postgresql-login! conn)
(guard (e (else #t)) (postgresql-execute-sql! conn "drop table test"))
(guard (e (else #t))
(postgresql-execute-sql! conn "create table test (data bytea)"))
(postgresql-terminate! conn)
;; let's do some benchmark
(postgresql-open-connection! conn)
(postgresql-login! conn)
(define data
(call-with-input-file "bench.scm" get-bytevector-all :transcoder #f))
(define (insert-it p)
(postgresql-bind-parameters! p data)
(postgresql-execute! p))
;; Re-using prepared statement
(let ((p (postgresql-prepared-statement
conn "insert into test (data) values ($1)")))
(time (dotimes (i 10) (insert-it p)))
(postgresql-close-prepared-statement! p))
(define (create-it)
(let ((p (postgresql-prepared-statement
conn "insert into test (data) values ($1)")))
(insert-it p)
(postgresql-close-prepared-statement! p)))
;; Creating prepared statement each time
(time (dotimes (i 10) (create-it)))
;; bye bye
(postgresql-terminate! conn)
I'm using (postgresql) library. <ad>BTW, I think this is the only portable library that can access database. So you gotta check it out. </ad> It's simply inserting the same binary data (in this case the script file itself) 10 times. One is re-using prepared statement, the other one is creating it each time. The difference is the following:
$ sash bench.scm ;; (dotimes (i 10) (insert-it p)) ;; 0.760319 real 0.008487 user 3.34e-40 sys ;; (dotimes (i 10) (create-it)) ;; 1.597769 real 0.014841 user 5.76e-40 sysMore than double. It's just doing 10 iterations but this much difference. (Please ignore the fact that the library itself is already slow.) There are probably couple of reasons including PostgreSQL itself but from the library perspective, sending messages to the server would be slow. Wherever a DB server is, even localhost, communication between script and the server is done via socket. And calling
postgresql-prepared-statement does at least 7 times of I/O (and 5 times for postgresql-close-prepared-statement!). So if I re-use it, then in total 120 times (12 x 10, of course) of I/O can be saved.Now, my ORM library hides low level operations such as creating prepared statement, connection management, etc. (that's what ORM should do, isn't it?). So keeping prepared statement in users' script isn't an option. Especially, there's no guarantee that users woudl get the same connection each time they do some operation. So it's better to manage it on the framework.
Sagittarius has
(cache lru) library, undocumented though, so first I thought I could use this. After modifying couple of lines and found out, no this isn't enough. The library only provides very simple cache mechanism. It even doesn't provide a way to get all objects inside the cache. It's okay if the object doesn't need any resource management, however prepared statements must be closed when it's no longer used. Plus, LRU may not be good enough for all situations so it might be better if users can specify which cache algorithm should be used.There are variety of cache algorithms. Implementing all of them would take a bit time. So it's better to make a framework or interface of cache. The framework/interface should have the following properties:
- Auto eviction and evict event handler
- Limitation of storage size (unlimited as well)
- A way to get all cached objects
- Implementation independent interface
2016-02-15
Server performance 2
So
Just giving up would be very easy way out but my consciousness doesn't allow me to do it (please let me go...). Thinking current HTTP server implementation uses 2 layers, Paella and Plato. The first one is the basic, then web framework. At least I can see which one would be slow. So I've just tried with bare Paella server. Copy&pasting the example and modify a bit like this:
Listing up what's actually done by server would help:
Then I've started doubting that the benchmark script itself is actually slow. I'm not sure how fast cURL itself is but forking it 1000 times and wait for them didn't sound fast. So I've written the following script:
2500 req/s isn't fast but for my purpose it's good enough for now. So I'll put this aside for now.
(net server) itself wasn't too bad performance, then there must be other culprit. To find out it, I usually use profiler however it can only work on single thread environment. That means it's impossible to use it on the server program written on top of (net server) library.Just giving up would be very easy way out but my consciousness doesn't allow me to do it (please let me go...). Thinking current HTTP server implementation uses 2 layers, Paella and Plato. The first one is the basic, then web framework. At least I can see which one would be slow. So I've just tried with bare Paella server. Copy&pasting the example and modify a bit like this:
(import (rnrs)
(net server)
(paella))
(define config (make-http-server-config :max-thread 10))
(define http-dispatcher
(make-http-server-dispatcher
(GET "/benchmark" (http-file-handler "index.html" "text/html"))))
(define server
(make-simple-server "8500" (http-server-handler http-dispatcher)
:config config))
(server-start! server)
Then uses the same script as before.The result is this:$ time ./benchmark.sh ./benchmark.sh 4.66s user 3.76s system 335% cpu 2.507 totalHmmm, bare server is already slow. So I can assume most of the time are consumed by the server, not the framework.
Listing up what's actually done by server would help:
- Converting socket to buffered port
- Parsing HTTP header
- Parsing request path.
- Parsing query string (if there is)
- Parsing mime (if there is)
- Parsing cookie (if there is)
- Calling handler
- Writing response
- Cleaning up
rfc5322-header-ref which is for referring header value called string-ci=? which calls string-foldcase internally. So changed it to call case folding once. And couple of more improvements. All of them, ideed, improved performance however calling header parser only 1000 times took 30ms from the beginning. So make it 15ms doesn't make that much change.Then I've started doubting that the benchmark script itself is actually slow. I'm not sure how fast cURL itself is but forking it 1000 times and wait for them didn't sound fast. So I've written the following script:
#!read-macro=sagittarius/bv-string
(import (rnrs)
(sagittarius socket)
(sagittarius control)
(time)
(util concurrent)
(getopt))
(define header
#*"GET /benchmark HTTP/1.1\r\n\
User-Agent: curl/7.35.0\r\n\
Host: localhost:8500\r\n\
Accept: */*\r\n\r\n")
(define (poke)
(define sock (make-client-socket "localhost" "8500"))
(socket-send sock header)
;; just poking
(socket-recv sock 256)
(socket-close sock))
(define (main args)
(with-args (cdr args)
((threads (#\t "threads") #t "10")
(unit (#\u "unit") #t "1000"))
(let* ((c (string->number unit))
(t (string->number threads))
(thread-pool (make-thread-pool t raise)))
(time (thread-pool-wait-all!
(dotimes (i (* c t) thread-pool)
(thread-pool-push-task! thread-pool poke))))
(thread-pool-release! thread-pool))))
Send fixed HTTP request and recieve the response (could be partially). -t option specifies how many threads should used and -u option specifies how many request should be done per thread. So if this ideed takes time, then my assumption is not correct. Lemme do it with bare HTTP server:$ sash bench.scm -t 100 -u 100 ;; (thread-pool-wait-all! (dotimes (i (* c t) thread-pool) (thread-pool-push-task! thread-pool poke))) ;; 4.052414 real 0.670089 user 1.255910 sys100 threads and 100 request per thread so in total 10000 request were send. Then it took 4 seconds, so 2500 req/s. It's faster than cURL version.
2500 req/s isn't fast but for my purpose it's good enough for now. So I'll put this aside for now.
2016-02-14
肉体改造部 第六週
今週(先週?)は風邪ひいてダウンしていた日が2日あったりした。月曜の夜にベッドの中で寒くて震えていたのはいい思い出である。普段は冬でも布団を蹴っ飛ばしてるのに・・・
計量結果:
計量結果:
- 体重: 72.8kg (-0.3kg)
- 体脂肪率: 23.6% (±0.0%)
- 筋肉率:42.6% (±0.0%)
2016-02-12
Server performance
Sagittarius has server framework library
I've created a very simple static page with Plato which is a web application framework bundled to Paella. It just return a HTML file. (although it does have some overhead...) It looks like this:
I don't have modern nice HTTP benchmark software like ApatchBench (because I'm lazy) so just used cURL and shell. The script looks like this:
The benchmark is done on default starting script which Plato generates. So number of threads are 10. Then this is the result:
If I run the above benchmark with 10 requests, then the result was like this:
Why it's so slow and gets slow when number of requests is increased? I think there are couple of reasons. The
Well in average it's 2.6sec per 1000 request so it is a bit faster like 300ms - 400ms. And using
(net server) and on top of this library I've written simple HTTP server and web framework, Paella. I don't use it in tight situation so performance isn't really matter for now. However if you write something you want to check how good or bad it is, don't you? And yes I've done simple benchmark and figured out it's horrible.I've created a very simple static page with Plato which is a web application framework bundled to Paella. It just return a HTML file. (although it does have some overhead...) It looks like this:
(library (plato webapp benchmark)
(export entry-point support-methods)
(import (rnrs) (paella) (plato) (util file))
(define (support-methods) '(GET))
(define (entry-point req)
(values 200 'file (build-path (plato-current-path (*plato-current-context*))
"index.html")
'("content-type" "text/html")))
)
The index.html has 200B data.I don't have modern nice HTTP benchmark software like ApatchBench (because I'm lazy) so just used cURL and shell. The script looks like this:
#!/bin/sh
invoke () {
curl http://localhost:8500/benchmark > /dev/null 2>&1
}
call () {
for i in `seq 1 1000`;
do
invoke &
done
}
call
wait
It's just create 1000 processes background and wait them.The benchmark is done on default starting script which Plato generates. So number of threads are 10. Then this is the result:
$ time ./benchmark.sh ./benchmark.sh 4.89s user 3.77s system 313% cpu 2.764 totalSo, I've done couple of times and average is approx 3 seconds per 1000 requests. So 300 Req/S. It's slow.
If I run the above benchmark with 10 requests, then the result was like this:
$ time ./benchmark.sh ./benchmark.sh 0.05s user 0.05s system 249% cpu 0.040 totalAnd 1 request is like this:
$ time ./benchmark.sh ./benchmark.sh 0.01s user 0.01s system 77% cpu 0.025 totalSo up to thread number, I can assume it does better, at least it's not increased 10 times. But if it's 100, then it's about 7 times more.
$ time ./benchmark.sh ./benchmark.sh 0.49s user 0.35s system 285% cpu 0.293 total1 to 10 is twice, but 10 to 100 is 7 times. Then 100 to 1000 is 10 times. Something isn't right to me.
Why it's so slow and gets slow when number of requests is increased? I think there are couple of reasons. The
(net server) uses combination of select (2)
and multithreading. When the server accepts the connection, then it
tries to find least used thread. After that it pushes the socket to the
found thread. The thread calls select if there's something
to read. Then invokes user defined procedure. After the invocation, it
checks if there's closed socket or not and waits input by select again. So flow is like this (n = number of thread, m = number of socket per thread):- Find least used thread. O(nm) (best case O(1) if none of the threads are used)
- Push socket to the thread. O(1)
- Handling request. O(m)
- Cleaning up sockets. O(m)
- Adding load balancing thread which simply manage priority queue
- Just asking the queue which thread is least loaded
- Code cleaning up
- Using
(util concurrent shared-queue)instead of manually managing sockets and locks - Don't assume write side shutdowned socket is not used.
- more...
$ time ./benchmark.sh ./benchmark.sh 4.61s user 3.76s system 317% cpu 2.633 totalYAHOOOOO!!!! 100ms faster!!! ... WHAAAATTTT!???
Well in average it's 2.6sec per 1000 request so it is a bit faster like 300ms - 400ms. And using
(util concurrent) made the server itself more robust (it sometimes hanged before). I think the server framework itself is not too bad but HTTP server. So that'd be the next step.
2016-02-07
肉体改造部 第五週
先週はなぜか書く機会を失った。
計量結果:
筋トレの負荷が足りない気がしているので、回数を倍にしているのだが、それでも足りない気がする(筋肉痛にすらならない)。ジムに行くべきなのだろうが、時間が取れないんだよなぁ。重り背負って腕立てとかかなぁ。
計量結果:
- 体重: 73.1kg (-0.6kg)
- 体脂肪率: 23.6% (-0.4%)
- 筋肉率:42.6% (+0.2%)
筋トレの負荷が足りない気がしているので、回数を倍にしているのだが、それでも足りない気がする(筋肉痛にすらならない)。ジムに行くべきなのだろうが、時間が取れないんだよなぁ。重り背負って腕立てとかかなぁ。
2016-01-29
Shadowing pattern identifier (macro bug)
I've been adding SRFIs, 57, 61, 87 and 99, to Sagittarius these days (means I just lost short term goal...). And found the bug (resolved). The bug was one of the longest lasting ones since I've re-written the macro expander. So it might be useful to share for someone who wants to write one macro expander from scratch.
The bug was about renaming pattern identifier. Well, more precisely, not renaming pattern identifier. On Sagittarius, pattern identifiers are preserved means they aren't renamed when
There would be 2 solution for this bug.
If it's just like this, probably I wouldn't write a blog post. The actual reason why is that how I found this bug. The bug was buried for a long time. I would say it's from the beginning (so ancient) but could also say since 0.5.0 (2 years). It's because the style of writing macros. I usually don't use the same name for pattern identifiers especially if ellipsis is involved. This is easier to debug. Now, when I was porting SRFI-57, I've noticed reference implementation was considerably slow. I think it's because the implementation uses macro generation macro generation macro generation... so on macro (not sure how much macro would be generated though). Unfortunately, Sagittarius' macro expander isn't so fast. Then I've found
Initially, I've thought this was implementation bug because the
What I want to say or put a note here is that bugs are found when you step out from your usual way. It'd be there, if I didn't port SRFI-57.It's worth to take a different way to do it.
The bug was about renaming pattern identifier. Well, more precisely, not renaming pattern identifier. On Sagittarius, pattern identifiers are preserved means they aren't renamed when
syntax-case is compiled. I don't remember why I needed to do this (should've written some comment but at that moment it was as clear as day, of course not anymore...) but if I change to rename it, then test cases, or even test itself, would run. Now, the bug was relating this non-renamed pattern identifier. You can see the issue but I also write the reproducible code here:
(import (rnrs))
(define-syntax foo
(lambda (x)
(define (derive-it k s)
(let loop ((r '()) (s s))
(syntax-case s ()
(() (datum->syntax k r))
(((f a b) rest ...)
(loop (cons #'(f a b) (cons #'(f a b) r)) #'(rest ...))))))
(syntax-case x ()
((k (foo . bar) ...)
(with-syntax ((((foo a b) ...)
(derive-it #'k #'((foo . bar) ...))))
#''((foo a b) ...))))))
(foo (f a b)) ;; -> shoulr return ((foo a b) (foo a b))
The problem is macro expander picked up not only foo of with-syntax but also foo of syntax-case. And bound input forms of these variables don't have the same length so macro expander signaled an error.There would be 2 solution for this bug.
- Rename pattern identifier (not even an option for me since it requires tacking macro expander again...)
- Consider shadowing of pattern identifiers.
with-syntax binds pattern identifiers and hides the same identifiers outside of the expression. So it can be considered creating a scope like let or so. Then taking only the top most binding (bound variables are like in environment, so I call it top) wouldn't be a problem, would it? So I took the easy path (#2).If it's just like this, probably I wouldn't write a blog post. The actual reason why is that how I found this bug. The bug was buried for a long time. I would say it's from the beginning (so ancient) but could also say since 0.5.0 (2 years). It's because the style of writing macros. I usually don't use the same name for pattern identifiers especially if ellipsis is involved. This is easier to debug. Now, when I was porting SRFI-57, I've noticed reference implementation was considerably slow. I think it's because the implementation uses macro generation macro generation macro generation... so on macro (not sure how much macro would be generated though). Unfortunately, Sagittarius' macro expander isn't so fast. Then I've found
syntax-case implementation on discussion archive. It was written for MzScheme but not so difficult to port for R6RS. Then faced the bug.Initially, I've thought this was implementation bug because the
syntax-case used at that moment could be different. So I've tested with other implementations and it worked. The error message said pretty close where it happened and dumped most of the information I needed, but just I couldn't see it in first glance. Maybe I don't want to see the fact that there's still macro related bugs...What I want to say or put a note here is that bugs are found when you step out from your usual way. It'd be there, if I didn't port SRFI-57.It's worth to take a different way to do it.
2016-01-26
syntax-case ポコ・ア・ポコ
syntax-caseを解説した日本語の文章というのは極端に少ないらしい。実際Googleで「syntax-case は」(日本語のみを検索する方法を知らないw)と検索しても、(オランダからだからかもしれないが)日本語の解説ページは自分が書いたものが一番上にヒットする。ここは一つ知ったかぶりをしてもう一つ検索結果を汚してもいいだろうと思ったので、syntax-caseの使い方を解説することにした(ここまで前置き)。想定読者はマクロはしってるけどSchemeのマクロはよく分からないという人としている。つまりsyntax-rulesを知らなくてもよい。ただ、マクロ自体は解説しないので、マクロとはなんぞやという疑問はこの記事を読んでも解決しないのであしからず。初めの一歩
Schemeの最新規格はR7RSだが、一つ前の規格R6RSで標準化されたsyntax-caseの使い方を解説する。まずは簡単な例を見てみよう。ここではwhenを定義することにする。
#!r6rs
(import (except (rnrs) when))
(define-syntax when
(lambda (x)
(syntax-case x ()
((_ test body1 body* ...)
#'(if test
(begin body1 body* ...))))))
(when 'a 'b) ;; -> b
(when #f #f) ;; -> unspecified
(when #f) ;; -> &syntax
最初のimportは知らなければおまじないと思ってくれればいい。次のwhenでマクロを定義している。syntax-caseは第一引数に構文オブジェクト、第二引数にリテラルリスト、それ以降にパターンと出力式のリストもしくは、パターン、フェンダー及び出力式のリストを受け取る。ここでは言葉を覚える必要はなく、そういうものだと思ってもらえればいい。
(define-syntax when
(lambda (x) ;; <- define-syntaxが受け取る手続きが受け取る引数が構文オブジェクト
(syntax-case x #| <- 第一引数:構文オブジェクト |# () #| >- 第二引数:リテラルリスト |#
;; パターンと出力式のリスト
;; (パターン 出力式)
;; もしくは
;; (パターン フェンダー 出力式) [フェンダーについては後述参照]
((_ test body1 body* ...)
#'(if test
(begin body1 body* ...))))))
出力式は基本構文オブジェクトを返す必要がある。#'を式につけると構文オブジェクトを返すようになる。#'はsyntaxの省略なので、上記のテンプレート部分は以下のようにも書ける:
(syntax (if test (begin body1 body* ...)))どちらを使うかは好みだが、筆者は
#'を使う方が見た目にも構文オブジェクトを返すことが分かりやすいのでこちらを使う。設問:
上記の
whenを参考にしてunlessを書いてみよ。ヒント:
unlessの展開形は(if (not test) (begin body1 body* ...)) のようになるはずである。パターンマッチ
パターンマッチとはなんぞやという人はあまりいないだろう。パターンを書くということは、入力式がそのパターンにマッチする必要がある。syntax-caseではリストもしくはベクタを入力式として分解することができる。基本的には識別子(*1)一つが要素一つに、可変長の入力を扱いたい場合は...(ellipsisと呼ばれる)を使う。例えば一つ以上の要素を持つリストにマッチさせるには以下のように書く:*1: コード上に現れるシンボルのこと。ここではクオートされたシンボルと区別するためにこう呼ぶ
(e e* ...) #| (1 2 3) ;; OK (これ以上でももちろんよい) (1) ;; OK () ;; NG |#一つのパターンに出てくる識別子は重複してはならないので、それぞれに別名をつける必要がある。筆者はよく可変長の入力にマッチするパターンの末尾に
*をつける。また、作法としてパターン識別子の先頭に?をつけてパターン識別子であることを分かりやすくするものもある。気をつけたいのは、...(ellipsis)は0個以上の入力にマッチするという点である。なので、以下のように書くと空リストにもマッチする:
(e* ...)ネストしたパターンを書くこともできる。例えば、要素の先頭がリスト(空リスト含む)であるリストにマッチするパターンは以下のように書ける:
((e1* ...) e2* ...) #| ((1 2 3) 4 5 6) ;; OK (() 4 5 6) ;; OK (()) ;; OK () ;; NG |#上記のパターンはリスト、ベクタ両方(ベクタの場合は
#をつけてベクタにする必要がある)に使える。ドット対のcdr部分にマッチさせることもできる。そのためには以下のように書く:
(a . d) #| (1 . 2) ;; OK (1 2 3) ;; OK (1 2 3) = (1 . (2 . (3 . ()))) (1) ;; OK () ;; NG |#ドット対のマッチと
...(ellipsis)を使うと、連想リストのキーと値にマッチすることも可能である。以下のように書く:
((a . d) ...) #| ((1 . 2)) ;; OK (1 2 3) ;; NG ((1 . 2) (3 . 4)) ;; OK () ;; OK |#組み合わせ次第で複雑な入力式にマッチさせることができる。
厳密な定義としてのパターンは以下のようになる:
- 識別子
- 定数 (文字、文字列、数値及びバイトベクタ)
(<パターン> ...)(<パターン> <パターン> ... . <パターン>)(<パターン> ... <パターン> <ellipsis> <パターン> ...)(<パターン> ... <パターン> <ellipsis> <パターン> ... . <パターン>)#(<パターン> ...)#(<パターン> ... <パターン> <ellipsis> <パターン> ...)
<パターン>は再帰的に定義されるので、リストパターンの中にベクタがあっても問題ない。また、定義で使われている...はパターンではなくパターンが複数あるという意味である。パターンの...は<ellipsis>となっているので注意されたい。ちなみに、パターンの定義は
syntax-rulesとほぼ同じなので、パターンマッチに関してはsyntax-caseを覚えればsyntax-rulesのものも使えるようになる(はずである)。出力式
出力式は基本的に構文オブジェクトを返す必要がある。大事なことなので二度目である。構文オブジェクトの生成にはsyntax (#')構文とquasisyntax (#`)構文の2種類ある。quasisyntaxはquasiquoteのsyntax版だと思えばよい。使い方は(あれば)次回やることにする(もしくはこちらを参照:Yet Another Syntax-case Explanation )。syntax構文が受け取る引数はテンプレートと呼ばる。テンプレートに指定できるのは、quoteと同じものが指定できる。quoteと違う点としてテンプレート内に現れてパターン変数(パターン内に現れた識別子のこと)がマッチした式に展開されるという点である。最初のwhenの例を見てみよう。whenのパターンは以下:
(_ test body1 body* ...)そして出力式は以下:
#'(if test (begin body1 body* ...)ここで、入力式として以下を受け取ったとしよう:
(when (zero? a) (display a) (newline) (do-with-a a))この入力式とパターンを対応づけると以下のようになる:
_ = when test = (zero? a) body1 = (display a) body* = ((newline) (do-with-a a))body*は
...(ellipsis)が付いているので可変長の入力を受け付けるが、ここでは便宜上複数要素を持つ一つのリストとしておく。パターンマッチでは言及していないが、_はプレースホルダーになるので、なんにでもマッチしかつ出力式では使用できないことに留意したい(*2)。*2: 要らない入力に名前を付けたくない場合に重宝する
ここまでくれば後は出力式に当てはめていくだけである。
...(ellipsis)を持つパターン変数はマッチした要素が一つずつ置換される、一つのリストではくなる、ので展開結果は以下のようになる。
(if (zero? a) (begin (zero? a) (display a) (newline) (do-with-a a)))とても簡単である。気をつけたい点としては
...(ellipsis)が付いているパターン変数はこれをつけないとマクロ展開器がエラーを投げることだろうか。マッチした入力式を展開する際は全ての入力式が消費される必要がある。もちろん、出力式に現れなかったパターン変数についてはその限りではない。フェンダー
用語として出してしまったので解説をしておく。フェンダーはパターンと出力式の間に入れることができるチェック用の式である。これが入っている場合はこの式が真の値を返した場合のみパターンにマッチしたと判定される。例えば以下のように使う:(define-syntax when
(lambda (x)
(syntax-case x ()
((_ test body1 body* ...)
(and (boolean? #'test) #'test) ;; testが#tであれば、if式は要らない
#'(begin body1 body* ...))
((_ test body1 body* ...)
#'(if test
(begin body1 body* ...))))))
フェンダー内で入力式を参照するにはsyntax構文を使ってパターン変数を展開してやる必要があることに注意したい。これ以外にも入力が識別子かどうか等のチェックなど用途はさまざまであるが、基本的にはパターンマッチ以外にチェックが必要な際に使う。ちなみに、パターンマッチは上から順に行われるため、上記のwhenの定義を逆にすると、フェンダーは評価されない。設問:
フェンダーを用いて
unlessを書いてみよ。長くなったのと「ポコ・ア・ポコ」なので今回はこれくらいにしておく。(要望があれば)次回は
syntax-caseが低レベル健全マクロと呼ばれる理由について書くことにする。
2016-01-24
肉体改造部 第三週
先週はルクセンブルグに行っていたので量れなかった。
計量結果:
二週間前に懸垂バーを買ったので、トレーニングに懸垂を追加している。合計で20回くらいしかやれないので(連続では最大で7回)、まぁいろいろ鈍っておる。導入初日は筋肉痛になったんだけど、2日目からはならない。やはり運動負荷が足りていないのだろうか?とりあえずは体重を落とす方を優先しているので、様子見かなぁ。
計量結果:
- 体重: 73.7kg (-0.4kg)
- 体脂肪率: 24.0% (+0.1%)
- 筋肉率:42.4% (+0.1%)
二週間前に懸垂バーを買ったので、トレーニングに懸垂を追加している。合計で20回くらいしかやれないので(連続では最大で7回)、まぁいろいろ鈍っておる。導入初日は筋肉痛になったんだけど、2日目からはならない。やはり運動負荷が足りていないのだろうか?とりあえずは体重を落とす方を優先しているので、様子見かなぁ。
2016-01-22
Dynamic compilation 2
Almost a year ago, I've wrote a post about compiling Scheme code to C. (See: Dynamic compilation (failure)) In the article, I've conclude that Sagittarius' VM is turned like crazy so ordinal C code wasn't match at all or something like that. After a year, I've noticed that there's a possibility that the compiler eliminated the expression and the VM just did some loop.
I'm not totally sure since when I've add code elimination and right now I don't have the version 0.6.2 (I think that's the version I've used) in my environment. So this might be totally bogus. Anyway, back then I used the following code:
Now, I might have some hope to turn up the VM. So prepare the shared object which provides
Even if this had siginificant improvement, I still need to resolve loads of things to native shared object from Scheme code. Such as:
It seems there's no easy way out for performance improvement. Maybe it's time to give up looking for this path.
I'm not totally sure since when I've add code elimination and right now I don't have the version 0.6.2 (I think that's the version I've used) in my environment. So this might be totally bogus. Anyway, back then I used the following code:
(define (fact n)
(let loop ((m 1) (r 1))
(if (= m n)
(* m r)
(loop (+ m 1) (* m r)))))
(time (dotimes (i 10000) (fact 1000)))
Now, the fact can be marked as transparent or no side effect because it seems it doesn't have any side effect nor consicing. Let me check.
(procedure-transparent? fact) ;; -> #tThe
procedure-transparent? is an internal procedure which is used by the compiler to eliminate dead code. So possibility is very high now. OK, let's check the VM instructions of the expression.
(disasm (lambda () (dotimes (i 10000) (fact 1000)))) ;; size: 15 ;; 0: CONSTI_PUSH(10000) ;; 1: CONSTI_PUSH(0) ;; 2: LREF_PUSH(1) ;; 3: LREF(0) ;; 4: BNGE 2 ;; 6: RET ;; 7: LREF_PUSH(0) ;; 8: LREF(1) ;; 9: ADDI(1) ;; 10: PUSH ;; 11: SHIFTJ(2 0) ;; 12: JUMP -11 ;; 14: RETBingo! There is no procedure call!
Now, I might have some hope to turn up the VM. So prepare the shared object which provides
fact. The C code is the following:#include <sagittarius.h>
#define LIBSAGITTARIUS_BODY
#include <sagittarius/extend.h>
static SgObject fact(SgObject *SG_FP, int SG_ARGC, void *data_)
{
SgObject m = SG_MAKE_INT(1), r = SG_MAKE_INT(1);
SgObject n = SG_FP[0];
while (TRUE) {
if (Sg_NumEq(m, SG_FP[0])) {
return Sg_Mul(m, r);
} else {
SgObject t1 = Sg_Add(m, SG_MAKE_INT(1));
SgObject t2 = Sg_Mul(m, r);
m = t1;
r = t2;
}
}
return SG_UNDEF; /* dummy */
}
static SG_DEFINE_SUBR(fact__STUB, 1, 0, fact, SG_FALSE, NULL);
SG_EXTENSION_ENTRY void CDECL Sg_Init_fact()
{
SgLibrary *lib = Sg_FindLibrary(SG_INTERN("(fact)"), TRUE);
SG_PROCEDURE_NAME(&fact__STUB) = SG_INTERN("fact");
Sg_InsertBinding(lib, SG_INTERN("fact"), &fact__STUB);
}
/*
gcc -lsagittarius fact.c -shared -o fact.so -fPIC -O3
*/
Now, benchmark. I've added set! to prevent the compiler optimisation.(load-dynamic-library "fact")
(import (time) (sagittarius control) (fact))
(define dummy)
;; Load C implementation first
(print fact)
(time (dotimes (i 10000) (set! dummy (fact 1000))))
;; Scheme implementation
(define (fact n)
(let loop ((m 1) (r 1))
(if (= m n)
(* m r)
(loop (+ m 1) (* m r)))))
(print fact)
(time (dotimes (i 10000) (set! dummy (fact 1000))))
#|
#<subr fact 1:0>
;; (dotimes (i 10000) (set! dummy (fact 1000)))
;; 6.604181 real 14.084392 user 0.174243 sys
#<closure fact 1:0>
;; (dotimes (i 10000) (set! dummy (fact 1000)))
;; 6.698120 real 14.32880 user 0.132117 sys
|#
Well, almost the same. C version is slightly faster. This is, I believe, because it doesn't have VM dispatch. But this is more or less error range.Even if this had siginificant improvement, I still need to resolve loads of things to native shared object from Scheme code. Such as:
- Mapping of Scheme procedure and C function
- Calling Scheme procedure from C
- Error handling especially unbound variables
- C compiler detection
- Etc. (macro, location so on)
call/cc friendly. There's a way to avoid some overhead, such as using CPS, however it'd be still the same or even slower then just running on VM (this is really proven by previous experience, unfortunately). On possible good future would be less consicing similar with the one mentioned unboxing in guile -- wingolog. Though, as long as I need to use CPS, it would most likely no more than trivial improvement.It seems there's no easy way out for performance improvement. Maybe it's time to give up looking for this path.
2016-01-10
肉体改造部 第一週
今週の結果
見た目の変化はない感じである。
負荷が足りてないのか、プロテインを飲んでいるからなのか分からないが筋肉痛にならなかったので、もう少し回数を増やすか何か背負って腕立てするかしてもいいかもしれない。
- 体重: 74.1kg (-0.6kg)
- 体脂肪率: 23.9% (-0.4%)
- 筋肉率: 42.3% (+ 0.1%)
見た目の変化はない感じである。
負荷が足りてないのか、プロテインを飲んでいるからなのか分からないが筋肉痛にならなかったので、もう少し回数を増やすか何か背負って腕立てするかしてもいいかもしれない。
Subscribe to:
Posts (Atom)