The full write-up is here. More context is given below.
The following exercise is from Tao’s lecture notes 7 from 247B. In the relevant part we consider,1 it states:
Proposition: Let \(1 < p < \infty\), and let \(b \in \mathcal{S}(\mathbb{R})\). Then the commutator of the Hilbert transform and the multiplier operator corresponding to \(b\) (i.e., \([A, B] = A B \ – B A\)) obeys an a priori bound $$\|[b, H] f\|_{L^p(\mathbb{R})} = \|b \cdot H f \ – H(b f)\|_{L^p(\mathbb{R})} \leq C_p \|b\|_{\text{BMO}(\mathbb{R})} \|f\|_{L^p(\mathbb{R})},$$ uniform over all \(f \in \mathcal{S}(\mathbb{R})\).
(This is part of a more general result of Coifman, Rochberg, and Weiss (1976),2 on the boundedness of \([b, T]\) on \(L^p(\mathbb{R}^d)\) for more general Calderón-Zygmund operators \(T\). This line of research has been further developed and generalized by M. Lacey and collaborators.)
The proof for this is pleasingly straightforward: one uses a paraproduct-like decomposition by taking the Littlewood-Paley expansion of \(b\), and multiplying by \(f\), splitting it as \(f_{\ll N}\), \(f_{\sim N}\), \(f_{\gg N}\) according to the index of the summand. One uses this to express the difference, and one can control the first two terms as paraproducts with a \(\text{BMO}\) function, and the theory of Carleson measures comes into play.
For the final contribution, we use duality, and see that there is an unexpected amount of cancellation: due to the local constancy of the multiplier for the Hilbert transform, all the terms with frequency support in far-off higher annuli will cancel and contribute nothing.
This is nice and straightforward, for the Schwartz class, where the paraproduct decomposition is generally technicality-free and perfectly legal (I assume this is why Tao chose to specify \(b \in \mathcal{S}(\mathbb{R})\) in the exercise). But the results of Coifman, Rochberg, Weiss indicate that this is true for more general functions than that; I sought to prove the \(\text{BMO}\) case, trying to reach it only with the standard dyadic discrete paraproduct machinery.
However, things become significantly more complicated when \(b \in \text{BMO}(\mathbb{R})\); \(b\) could be unbounded, for instance, and even if it is not, the Littlewood-Paley decomposition that powers the paraproduct generally has issues on \(L^{\infty}\) functions; it works best for \(L^p\), \(1 < p < \infty\).
The extension from Schwartz \(b\) to general \(b\) required a bit of thought: this led to a week of me searching on Google for things like “boundedness of smooth spatial cutoff multipliers against \(\text{BMO}\) functions,” trying to find a soft argument that would allow me to express \(b\) as a bounded weak limit of compactly-supported (or Schwartz) functions, weakly in \(\text{BMO}\). Embarrassingly (it took me several days, after I began searching, to realize this), I saw that the fix is quite simple: even though the identity \(b = \sum_N b_N\) is generally invalid for \(b \in L^{\infty}(\mathbb{R})\), when passing to everything in a weak sense, this holds “up to constants”: the low-frequency term \(P_{\leq N} b\) (perhaps after passing to a subsequence) tends to some constant in the weak\(^*\) topology on \(L^{\infty}(\mathbb{R})\), as \(N \downarrow 0\). Thanks to the commutator structure, this constant correction gets subtracted away; it vanishes!
This allows us to proceed as before; once again expanding in paraproduct-like sums, though being careful to separate the very-low-frequency piece for separate consideration, and taking various limits, one gets the desired bound via duality: $$|\langle [b, H] f, g \rangle| \lesssim_p \|b\|_{\text{BMO}(\mathbb{R})} \|f\|_{L^p(\mathbb{R})} \|g\|_{L^{p’}(\mathbb{R})},$$ for \(b \in L^{\infty}(\mathbb{R})\), and Schwartz \(f\) and \(g\) satisfying various technical assumptions that further pave the way for an easy application of a paraproduct-like decomposition. After rearranging, and using more soft arguments and checking convergence carefully, one gets the following, which is the main highlight of the write-up:
Theorem: Let \(1 < p < \infty\), and fix \(b \in \text{BMO}(\mathbb{R})\). Then \([b, H] f \in L^p(\mathbb{R})\) for all \(f \in \mathcal{S}(\mathbb{R})\), and we have $$\|[b, H] f\|_{L^p(\mathbb{R})} \lesssim_p \|b\|_{\text{BMO}(\mathbb{R})} \|f\|_{L^p(\mathbb{R})}.$$
In other words, the paraproduct proof does still work to give the result for \(\text{BMO}\) functions, not just Schwartz ones.
This is also somewhat surprising, for the following reason: for \(f \in \mathcal{S}(\mathbb{R})\), it is certainly true that \(b f \in L^p(\mathbb{R})\) for any \(1 \leq p < \infty\), though we may not have any control of the \(L^p\) norm of \(b f\) in terms of \(\|f\|_{L^p(\mathbb{R})}\) alone. But the commutator estimate then makes clear that \(b \cdot H f\) is also an \(L^p\) function. This is surprising to me, because \(H f\) for \(f \in \mathcal{S}(\mathbb{R})\) not vanishing at the frequency origin does not have to be Schwartz; the Hilbert transform is smooth, but does not need to have the nice decay of a Schwartz function. As a consequence, it is rather interesting (and, to me, not a priori obvious) that for any \(f \in \mathcal{S}(\mathbb{R})\), \(b \cdot H f\) still is in \(L^p\) for any \(1 < p < \infty\). Even if the dependence in norm is unbounded, it’s interesting that the image of \(\mathcal{S}(\mathbb{R})\) under \(M_b H\) still never leaves any \(L^p\).
Work begun: August 2, 2025. Finished: August 2, 2025. (Part I.)
Work begun: August 19, 2025. Finished: August 19, 2025. (Part II.)
- I haven’t yet figured out how to complete the other half of this exercise: that there is a double-sided numerical equivalence between the operator norm of the commutator and the \(\text{BMO}\) norm of the multiplier. If anyone knows how to implement the hint in the exercise, that one should consider \(\langle [b, H] f, g \rangle\) for \(f, g\) supported in nearby intervals of the same size, I’d certainly be glad to hear about it; due to the difficulty in computing Hilbert transforms in closed form even for well-known elementary functions, much less ascertaining pointwise behavior of \(H(b f)\) for an unknown \(b\), I certainly have no idea how one is supposed to do this. ↩︎
- R. R. Coifman, R. Rochberg & G. Weiss, Factorization Theorems for Hardy Spaces in Several Variables, 103 Ann. Math. 611 (1976). doi:10.2307/1970954. MR 0412721. ↩︎
Leave a Reply