SearcharxivSearch

arXiv subjects

Mykola Pinchuk

Publications and source records attributed to Mykola Pinchuk.

8 recordsLinked to original sources

TML-Bench: Benchmark for Data Science Agents on Tabular ML Tasks

Autonomous coding agents can produce strong tabular baselines quickly on Kaggle-style tasks. Practical value depends on end-to-end correctness and reliability under time limits. This paper introduces TML-Bench, a tabular benchmark for data science agents on Kaggle-style tasks. This paper evaluates 10 OSS LLMs on four Kaggle competitions and three time budgets (240s, 600s, and 1200s). Each model is run five times per task and budget. A run is successful if it produces a valid submission and a private-holdout score on hidden labels that are not accessible to the agent. This paper reports median performance, success rates, and run-to-run variability. MiniMax-M2.1 model achieves the best aggregate performance score on all four competitions under the paper's primary aggregation. Average performance improves with larger time budgets. Scaling is noisy for some individual models at the current run count. Code and materials are available at https://github.com/MykolaPinchuk/TML-bench/tree/master.

cs.LG

Time Aggregation Features for XGBoost Models

This paper studies time aggregation features for XGBoost models in click-through rate prediction. The setting is the Avazu click-through rate prediction dataset with strict out-of-time splits and a no-lookahead feature constraint. Features for hour H use only impressions from hours strictly before H. This paper compares a strong time-aware target encoding baseline to models augmented with entity history time aggregation under several window designs. Across two rolling-tail folds on a deterministic ten percent sample, a trailing window specification improves ROC AUC by about 0.0066 to 0.0082 and PR AUC by about 0.0084 to 0.0094 relative to target encoding alone. Within the time aggregation design grid, event count windows provide the only consistent improvement over trailing windows, and the gain is small. Gap windows and bucketized windows underperform simple trailing windows in this dataset and protocol. These results support a practical default of trailing windows, with an optional event count window when marginal ROC AUC gains matter.

cs.LG

Intra-tree Column Subsampling Hinders XGBoost Learning of Ratio-like Interactions

Many applied problems contain signal that becomes clear only after combining multiple raw measurements. Ratios and rates are common examples. In gradient boosted trees, this combination is not an explicit operation: the model must synthesize it through coordinated splits on the component features. We study whether intra-tree column subsampling in XGBoost makes that synthesis harder. We use two synthetic data generating processes with cancellation-style structure. In both, two primitive features share a strong nuisance factor, while the target depends on a smaller differential factor. A log ratio cancels the nuisance and isolates the signal. We vary colsample_bylevel and colsample_bynode over s in {0.4, 0.6, 0.8, 0.9}, emphasizing mild subsampling (s >= 0.8). A control feature set includes the engineered ratio, removing the need for synthesis. Across both processes, intra-tree column subsampling reduces test PR-AUC in the primitives-only setting. In the main process the relative decrease reaches 54 percent when both parameters are set to 0.4. The effect largely disappears when the engineered ratio is present. A path-based co-usage metric drops in the same cells where performance deteriorates. Practically, if ratio-like structure is plausible, either avoid intra-tree subsampling or include the intended ratio features.

cs.LG

Zero-Leverage Puzzle

In this paper, I examine why some firms have zero leverage. I fail to find evidence that firms are unlevered because of managerial entrenchment since these firms do not have weaker corporate governance. I reject the hypothesis that firms become zero-leverage after prolonged periods of high market valuation, since before levering these firms do not suffer from declining valuations and continue to issue large amounts of equity. I find strong evidence in favor of the financial constraints explanation of the zero-leverage puzzle. Zero-leverage firms appear to be financially constrained using three different measures of financial constraints. I obtain mixed evidence on the financial flexibility hypothesis since all-equity firms increase investments and acquisitions after levering, but the probability of their levering decreased during the financial crisis. My results suggest that financial constraints are the first-order the driver of zero-leverage behavior and are more important than less obvious explanations such as managerial entrenchment.

q-fin.GN

Customer Momentum

This paper examines customer momentum, defined as a positive relationship between a firm's returns and past returns of its customers. I confirm previous evidence (Cohen and Frazzini 2008) that customer momentum is both statistically and economically significant. Long-short equally-weighted (value-weighted) decile portfolio generates a monthly return of 122 (106) basis points and a t-statistic above 4 (2.8) with respect to Fama-French factor models. The paper reports that customer momentum neither explains nor is explained by price momentum and earnings momentum. Customer momentum is partially driven by the lead-lag relationship between small and large stocks. I find that in the post-discovery sample, customer momentum has a smaller magnitude and loses statistical significance. The results are consistent with the hypothesis that after its discovery, customer momentum decreased due to exploitation by investors.

q-fin.PR

Bitcoin Does Not Hedge Inflation

This paper examines the response of major cryptocurrencies to macroeconomic news announcements (MNA). While other cryptocurrencies exhibit no reaction to major MNA, Bitcoin responds negatively to inflation surprise. Price of Bitcoin decreases by 24 bps in response to a 1 standard deviation inflationary surprise. This reaction is inconsistent with widely-held beliefs of practitioners that Bitcoin can hedge inflation. I do not find support for the hypothesis that the negative response of Bitcoin to inflation is due to its negative exposure to interest rates. Instead, I find support for the hypothesis that Bitcoin is strongly affected by the shift in consumption-savings decisions, driven by the rise in inflation. Consistent with this view, Bitcoin has negative exposure to a proxy for the consumption-savings ratio.

q-fin.PR

Labor Income Risk and the Cross-Section of Expected Returns

This paper explores asset pricing implications of unemployment risk from sectoral shifts. I proxy for this risk using cross-industry dispersion (CID), defined as a mean absolute deviation of returns of 49 industry portfolios. CID peaks during periods of accelerated sectoral reallocation and heightened uncertainty. I find that expected stock returns are related cross-sectionally to the sensitivities of returns to innovations in CID. Annualized returns of the stocks with high sensitivity to CID are 5.9% lower than the returns of the stocks with low sensitivity. Abnormal returns with respect to the best factor model are 3.5%, suggesting that common factors can not explain this return spread. Stocks with high sensitivity to CID are likely to be the stocks, which benefited from sectoral shifts. CID positively predicts unemployment through its long-term component, consistent with the hypothesis that CID is a proxy for unemployment risk from sectoral shifts.

q-fin.PR

Monetary Uncertainty as a Determinant of the Response of Stock Market to Macroeconomic News

This paper examines the effect of macroeconomic news announcements (MNA) on the stock market. Stocks exhibit a strong positive response to major MNA: 1 standard deviation of MNA surprise causes 11-25 bps higher returns. This response is highly time-varying and is weaker during periods of high monetary uncertainty. I decompose this response into cash flow and risk-free rate channels. 1 standard deviation of good MNA surprise leads to plus 30 bps returns from the cash flow channel and minus 23 bps per 1\% of monetary uncertainty from the risk-free rate channel. Risk-free rate channel is time-varying and is stronger when monetary uncertainty is high. High levels of monetary uncertainty mask the strong positive response of stocks to MNA, which explains why past research failed to detect this relation.

q-fin.PR